Towards AITowards AIToneBench

Kimi K3 vs Grok 4.6 (high)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.3 to 88.1.

Kimi K3
#8 Elo 2240 · 89.3/100
Grok 4.6 (high)
#18 Elo 2144 · 88.1/100
Cost / script
$0.266 vs $0.115
Human baseline
90.3 Kimi K3 falls below it · Grok 4.6 (high) falls below it

The verdict

Kimi K3 takes this one, and the margin is now clean. Kimi sits seventh at 2240.5, Grok 4.6 seventeenth at 2144.5, a 96-point gap, and this time the confidence intervals no longer overlap, so treat the lead as real. On raw scores it is 89.26 against 88.14, a shade over a point. The separation still comes from craft: Kimi leads continuity and emotion 88.33 to 86.27 and cue quality 87.41 to 86.43. Grok answers on numbers, 88.24 to 87.31 on the anti-slop metric, while hooks are a wash, 90.07 to 90.03. Then the ledger flips. Grok costs about twelve cents per script to Kimi's about twenty-seven cents, both exact figures, so the cheaper model here is the closed one. Kimi's counter is that it is open weights, which matters if you ever want to self-host. Pay a bit over 2x for cleaner scene-to-scene flow, or take the savings and fix the transitions yourself.

Pick Kimi K3 if you want the stronger scene-to-scene continuity and cue work, value open weights you can self-host, and can absorb a bit over 2x the per-script cost.
Pick Grok 4.6 (high) if you want most of the quality at about twelve cents per script and cleaner handling of numbers matters more to you than seamless transitions.

Metric by metric

Blue bars: Kimi K3. Orange bars: Grok 4.6 (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.5
88.8
Writing Craft & Clarity13% weight
89.4
88.0
Substance, Accuracy & Value15% weight
89.5
88.0
Continuity & Emotion14% weight
88.3
86.3
YouTube Best Practices12% weight
88.7
87.1
Hook Strength10% weight
90.1
90.0
Length Adherence8% weight
89.0
88.3
Slop Score (EQ-Bench + ours)5% weight
91.9
91.8
Visual Cue Quality4% weight
87.4
86.4

Everything else that differs

Kimi K3Grok 4.6 (high)
Overall / 10089.388.1
Writing Elo22402144
Run-to-run spread (± overall std)1.5201.430
Cost per script (USD)0.2660.115
Avg latency (s)372.1199.2
Open weightsYesNo

Full scorecards: Kimi K3 · Grok 4.6 (high). How scoring works: methodology.

← All comparisons