Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 85.8.
This one isn't close. Kimi K3 wins at 89.3 overall against Grok 4.5's 85.8, and the Elo intervals don't even touch, so no hedging needed. Credit where it's due first: Grok is fast, under 30 seconds per script, and at roughly 4 cents it costs about a fifth of Kimi. That combination is genuinely useful when you're iterating. But as a writer? Grok's biggest problem is discipline. It scores 78.7 on length adherence versus Kimi's 91.1, which in practice means scripts that overshoot or undershoot the target and need trimming by hand. Continuity and emotion tell the same story, 83.5 against 88.3: sections that sit next to each other instead of flowing into each other. Kimi leads on nearly every metric while staying open weights, which I appreciate, and its hooks at 90.2 are the best in this pair. If the script is going anywhere near publish, Kimi. If you're testing ten hook ideas in a minute, Grok earns its spot.
Pick Kimi K3 if the script is meant to publish, since it wins nearly every metric and stays open weights.
Pick Grok 4.5 if you're iterating fast on cheap drafts and will fix length and flow yourself.
Blue bars: Kimi K3 (thinking). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 (thinking) | Grok 4.5 | |
|---|---|---|
| Overall / 100 | 89.2 | 85.8 |
| Writing Elo | 2416 | 2042 |
| Run-to-run spread (± overall std) | 1.650 | 2.810 |
| Cost per script (USD) | 0.187 | 0.038 |
| Avg latency (s) | 267.2 | 41.7 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 (thinking) · Grok 4.5. How scoring works: methodology.
← All comparisons