Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 88.2 to 86.6.
Kimi K3 wins this one, and the intervals say it's real: Kimi sits at #11 on the board with 2180.2 Elo, Grok 4.6 at #26 with 2035.3, and their confidence intervals don't overlap. Overall it's 88.19 to 86.61, a 1.58-point gap. Most of that gap is the hook: 90.19 against 80.81, by far the widest split on the card. Kimi also takes writing quality, 88.45 vs 86.49, and length adherence, 90.07 to 88.03, while substance is a wash at 87.57 to 87.53. Grok fights back on production details, winning visual cues 86.63 vs 84.6 and the numbers-and-slop metric 87.01 to 86.27. The records lean the same way: 1041 wins and 6 losses for Kimi against Grok's 941 and 99. The bill barely separates them, about twenty-six cents per script for Kimi against about twenty cents for Grok, and Kimi is open weights on top. Saving six cents doesn't buy back the hook.
Pick Kimi K3 if you want open weights and the clearly stronger opening, a 90.19 hook against Grok's 80.81, plus better writing quality, at about twenty-six cents a script.
Pick Grok 4.6 if you want slightly cheaper scripts at about twenty cents each with better visual cues and cleaner handling of numbers, and can live with the weaker hooks.
Blue bars: Kimi K3. Orange bars: Grok 4.6. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | Grok 4.6 | |
|---|---|---|
| Overall / 100 | 88.2 | 86.6 |
| Writing Elo | 2180 | 2035 |
| Run-to-run spread (± overall std) | 1.680 | 2.530 |
| Cost per script (USD) | 0.260 | 0.204 |
| Avg latency (s) | 234.3 | 243.1 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 · Grok 4.6. How scoring works: methodology.
← All comparisons