Towards AITowards AIToneBench

Kimi K3 vs GPT-5.6 Sol (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 88.4.

Kimi K3
#8 Elo 2438 · 89.4/100
GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
Cost / script
$0.263 vs $0.163
Human baseline
92.7 Kimi K3 falls below it · GPT-5.6 Sol (high) falls below it

The verdict

Best open-weights model on the board against the best non-Anthropic closed one. Kimi K3 takes it at #8 to GPT's #17, 2437.5 Elo against 2315.4, intervals clear, and 89.39 overall to 88.09. But that summary hides a real split. GPT wins the two mechanical metrics decisively: anti-slop and numbers at 90.60 against 87.36, and cue quality at 91.02 against 88.30. Those are both about three points, and both are things you'd otherwise fix by hand. Kimi wins everything about the writing itself. Tone and voice 89.61 to 87.31, continuity 88.28 to 86.07, hooks 89.88 to 87.51. So the choice is genuinely about what you value: GPT gives you cleaner mechanics, Kimi gives you a better piece of writing. GPT is much faster, 106 seconds against 367, and cheaper, 16 cents against 25. Kimi's argument is open weights plus the better read.

Pick Kimi K3 if the writing itself is what matters, voice and flow and openings, and you want open weights.
Pick GPT-5.6 Sol (high) if you want the cleanest cues and number handling, back in under two minutes, for less money.

Metric by metric

Blue bars: Kimi K3. Orange bars: GPT-5.6 Sol (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.6
87.3
Writing Craft & Clarity13% weight
89.3
88.1
Substance, Accuracy & Value15% weight
89.5
90.1
Continuity & Emotion14% weight
88.3
86.1
YouTube Best Practices12% weight
88.4
86.5
Hook Strength10% weight
89.9
87.5
Length Adherence8% weight
90.5
91.7
Slop Score (EQ-Bench + ours)5% weight
92.0
93.4
Visual Cue Quality4% weight
88.3
91.0

Everything else that differs

Kimi K3GPT-5.6 Sol (high)
Overall / 10089.488.4
Writing Elo24382315
Run-to-run spread (± overall std)1.5701.750
Cost per script (USD)0.2630.163
Avg latency (s)367.1106.4
Open weightsYesNo

Full scorecards: Kimi K3 · GPT-5.6 Sol (high). How scoring works: methodology.

← All comparisons