Towards AITowards AIToneBench

GPT-5.6 Sol (ultra) vs Kimi K3 (thinking)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 88.4. Their confidence intervals overlap, so treat the order as close rather than settled.

GPT-5.6 Sol (ultra)
#16 Elo 2316 · 88.4/100
Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
Cost / script
$0.137 vs $0.187
Human baseline
92.7 GPT-5.6 Sol (ultra) falls below it · Kimi K3 (thinking) falls below it

The verdict

Best open-weights model against GPT's new top rung. Kimi K3 thinking takes the board at #12 to #16, 89.25 overall against 88.40, but under a point separates them and this one is genuinely close. GPT wins the mechanical metrics it always wins: length adherence 94.34 against 91.10, cue quality 91.17 against 87.67, anti-slop 90.60 against 87.24. Kimi wins the writing: continuity 88.28 against 85.72, hooks 90.18 against 86.39, tone 89.31 against 86.99. Same split as the rest of the GPT-versus-Kimi story, just tighter now that ultra exists. Practical differences: Kimi costs 19 cents against GPT's 16, and both are slow, 267 seconds against 421. So GPT is cheaper and mechanically cleaner, Kimi is open and writes better. Both sit under the 92.74 human baseline.

Pick GPT-5.6 Sol (ultra) for the tightest length control and cleanest cues, at two thirds the price.
Pick Kimi K3 (thinking) if you want open weights and the better piece of writing, especially the hooks and flow.

Metric by metric

Blue bars: GPT-5.6 Sol (ultra). Orange bars: Kimi K3 (thinking). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.0
89.3
Writing Craft & Clarity13% weight
88.1
89.3
Substance, Accuracy & Value15% weight
90.5
89.2
Continuity & Emotion14% weight
85.7
88.3
YouTube Best Practices12% weight
86.1
87.8
Hook Strength10% weight
86.4
90.2
Length Adherence8% weight
94.3
91.1
Slop Score (EQ-Bench + ours)5% weight
93.5
91.6
Visual Cue Quality4% weight
91.2
87.7

Everything else that differs

GPT-5.6 Sol (ultra)Kimi K3 (thinking)
Overall / 10088.489.2
Writing Elo23162416
Run-to-run spread (± overall std)1.9701.650
Cost per script (USD)0.1370.187
Avg latency (s)420.9267.2
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (ultra) · Kimi K3 (thinking). How scoring works: methodology.

← All comparisons