Towards AITowards AIToneBench

Kimi K3 (thinking) vs GPT-5.6 Sol (ultra)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 88.4. Their confidence intervals overlap, so treat the order as close rather than settled.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
GPT-5.6 Sol (ultra)
#16 Elo 2316 · 88.4/100
Cost / script
$0.187 vs $0.137
Human baseline
92.7 Kimi K3 (thinking) falls below it · GPT-5.6 Sol (ultra) falls below it

The verdict

Best open-weights model against GPT's new top rung. Kimi K3 thinking takes the board at #12 to #16, 89.25 overall against 88.40, but under a point separates them and this one is genuinely close. GPT wins the mechanical metrics it always wins: length adherence 94.34 against 91.10, cue quality 91.17 against 87.67, anti-slop 90.60 against 87.24. Kimi wins the writing: continuity 88.28 against 85.72, hooks 90.18 against 86.39, tone 89.31 against 86.99. Same split as the rest of the GPT-versus-Kimi story, just tighter now that ultra exists. Practical differences: Kimi costs 19 cents against GPT's 16, and both are slow, 267 seconds against 421. So GPT is cheaper and mechanically cleaner, Kimi is open and writes better. Both sit under the 92.74 human baseline.

Pick GPT-5.6 Sol (ultra) for the tightest length control and cleanest cues, at two thirds the price.
Pick Kimi K3 (thinking) if you want open weights and the better piece of writing, especially the hooks and flow.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
87.0
Writing Craft & Clarity13% weight
89.3
88.1
Substance, Accuracy & Value15% weight
89.2
90.5
Continuity & Emotion14% weight
88.3
85.7
YouTube Best Practices12% weight
87.8
86.1
Hook Strength10% weight
90.2
86.4
Length Adherence8% weight
91.1
94.3
Slop Score (EQ-Bench + ours)5% weight
91.6
93.5
Visual Cue Quality4% weight
87.7
91.2

Everything else that differs

Kimi K3 (thinking)GPT-5.6 Sol (ultra)
Overall / 10089.288.4
Writing Elo24162316
Run-to-run spread (± overall std)1.6501.970
Cost per script (USD)0.1870.137
Avg latency (s)267.2420.9
Open weightsYesNo

Full scorecards: Kimi K3 (thinking) · GPT-5.6 Sol (ultra). How scoring works: methodology.

← All comparisons