Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 85.8.
Kimi K3 at #8 against Grok 4.5 at #30, 89.39 overall to 85.8. Four and a quarter points, intervals well clear. The decisive metric is length adherence: Kimi 90.46, Grok 78.67. Fourteen points. Grok does not hit a target word count, and on long-form that alone justifies the ranking gap. Continuity is the other big one, 88.28 to 83.50. Where Grok stays competitive is substance at 87.42 against 89.6, and hooks at 87.96 against 89.36, so the thinking and the openings are not the problem. It's the shape of the finished piece. Now the other direction: Grok costs under four cents a script and returns in 42 seconds. Kimi costs 26 cents and takes over six minutes. That's seven times the price and ten times the wait for four points. If you're drafting at volume and editing anyway, that math favours Grok more than the leaderboard position suggests.
Pick Kimi K3 if the draft needs to hold its length and flow, and you want open weights.
Pick Grok 4.5 if throughput is the constraint. Seven times cheaper, ten times faster, and the ideas mostly land.
Blue bars: Kimi K3. Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | Grok 4.5 | |
|---|---|---|
| Overall / 100 | 89.4 | 85.8 |
| Writing Elo | 2438 | 2042 |
| Run-to-run spread (± overall std) | 1.570 | 2.810 |
| Cost per script (USD) | 0.263 | 0.038 |
| Avg latency (s) | 367.1 | 41.7 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 · Grok 4.5. How scoring works: methodology.
← All comparisons