Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.3 to 88.1.
Kimi K3 takes this one, and the margin is now clean. Kimi sits seventh at 2240.5, Grok 4.6 seventeenth at 2144.5, a 96-point gap, and this time the confidence intervals no longer overlap, so treat the lead as real. On raw scores it is 89.26 against 88.14, a shade over a point. The separation still comes from craft: Kimi leads continuity and emotion 88.33 to 86.27 and cue quality 87.41 to 86.43. Grok answers on numbers, 88.24 to 87.31 on the anti-slop metric, while hooks are a wash, 90.07 to 90.03. Then the ledger flips. Grok costs about twelve cents per script to Kimi's about twenty-seven cents, both exact figures, so the cheaper model here is the closed one. Kimi's counter is that it is open weights, which matters if you ever want to self-host. Pay a bit over 2x for cleaner scene-to-scene flow, or take the savings and fix the transitions yourself.
Pick Kimi K3 if you want the stronger scene-to-scene continuity and cue work, value open weights you can self-host, and can absorb a bit over 2x the per-script cost.
Pick Grok 4.6 (high) if you want most of the quality at about twelve cents per script and cleaner handling of numbers matters more to you than seamless transitions.
Blue bars: Kimi K3. Orange bars: Grok 4.6 (high). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | Grok 4.6 (high) | |
|---|---|---|
| Overall / 100 | 89.3 | 88.1 |
| Writing Elo | 2240 | 2144 |
| Run-to-run spread (± overall std) | 1.520 | 1.430 |
| Cost per script (USD) | 0.266 | 0.115 |
| Avg latency (s) | 372.1 | 199.2 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 · Grok 4.6 (high). How scoring works: methodology.
← All comparisons