Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 88.4. Their confidence intervals overlap, so treat the order as close rather than settled.
Best open-weights model against GPT's new top rung. Kimi K3 thinking takes the board at #12 to #16, 89.25 overall against 88.40, but under a point separates them and this one is genuinely close. GPT wins the mechanical metrics it always wins: length adherence 94.34 against 91.10, cue quality 91.17 against 87.67, anti-slop 90.60 against 87.24. Kimi wins the writing: continuity 88.28 against 85.72, hooks 90.18 against 86.39, tone 89.31 against 86.99. Same split as the rest of the GPT-versus-Kimi story, just tighter now that ultra exists. Practical differences: Kimi costs 19 cents against GPT's 16, and both are slow, 267 seconds against 421. So GPT is cheaper and mechanically cleaner, Kimi is open and writes better. Both sit under the 92.74 human baseline.
Pick GPT-5.6 Sol (ultra) for the tightest length control and cleanest cues, at two thirds the price.
Pick Kimi K3 (thinking) if you want open weights and the better piece of writing, especially the hooks and flow.
Blue bars: GPT-5.6 Sol (ultra). Orange bars: Kimi K3 (thinking). Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (ultra) | Kimi K3 (thinking) | |
|---|---|---|
| Overall / 100 | 88.4 | 89.2 |
| Writing Elo | 2316 | 2416 |
| Run-to-run spread (± overall std) | 1.970 | 1.650 |
| Cost per script (USD) | 0.137 | 0.187 |
| Avg latency (s) | 420.9 | 267.2 |
| Open weights | No | Yes |
Full scorecards: GPT-5.6 Sol (ultra) · Kimi K3 (thinking). How scoring works: methodology.
← All comparisons