Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 88.4.
Best open-weights model on the board against the best non-Anthropic closed one. Kimi K3 takes it at #8 to GPT's #17, 2437.5 Elo against 2315.4, intervals clear, and 89.39 overall to 88.09. But that summary hides a real split. GPT wins the two mechanical metrics decisively: anti-slop and numbers at 90.60 against 87.36, and cue quality at 91.02 against 88.30. Those are both about three points, and both are things you'd otherwise fix by hand. Kimi wins everything about the writing itself. Tone and voice 89.61 to 87.31, continuity 88.28 to 86.07, hooks 89.88 to 87.51. So the choice is genuinely about what you value: GPT gives you cleaner mechanics, Kimi gives you a better piece of writing. GPT is much faster, 106 seconds against 367, and cheaper, 16 cents against 25. Kimi's argument is open weights plus the better read.
Pick Kimi K3 if the writing itself is what matters, voice and flow and openings, and you want open weights.
Pick GPT-5.6 Sol (high) if you want the cleanest cues and number handling, back in under two minutes, for less money.
Blue bars: Kimi K3. Orange bars: GPT-5.6 Sol (high). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | GPT-5.6 Sol (high) | |
|---|---|---|
| Overall / 100 | 89.4 | 88.4 |
| Writing Elo | 2438 | 2315 |
| Run-to-run spread (± overall std) | 1.570 | 1.750 |
| Cost per script (USD) | 0.263 | 0.163 |
| Avg latency (s) | 367.1 | 106.4 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 · GPT-5.6 Sol (high). How scoring works: methodology.
← All comparisons