Towards AITowards AIToneBench

Kimi K3 vs Grok 4.6

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 88.2 to 86.6.

Kimi K3
#11 Elo 2180 · 88.2/100
Grok 4.6
#26 Elo 2035 · 86.6/100
Cost / script
$0.260 vs $0.204
Human baseline
90.2 Kimi K3 falls below it · Grok 4.6 falls below it

The verdict

Kimi K3 wins this one, and the intervals say it's real: Kimi sits at #11 on the board with 2180.2 Elo, Grok 4.6 at #26 with 2035.3, and their confidence intervals don't overlap. Overall it's 88.19 to 86.61, a 1.58-point gap. Most of that gap is the hook: 90.19 against 80.81, by far the widest split on the card. Kimi also takes writing quality, 88.45 vs 86.49, and length adherence, 90.07 to 88.03, while substance is a wash at 87.57 to 87.53. Grok fights back on production details, winning visual cues 86.63 vs 84.6 and the numbers-and-slop metric 87.01 to 86.27. The records lean the same way: 1041 wins and 6 losses for Kimi against Grok's 941 and 99. The bill barely separates them, about twenty-six cents per script for Kimi against about twenty cents for Grok, and Kimi is open weights on top. Saving six cents doesn't buy back the hook.

Pick Kimi K3 if you want open weights and the clearly stronger opening, a 90.19 hook against Grok's 80.81, plus better writing quality, at about twenty-six cents a script.
Pick Grok 4.6 if you want slightly cheaper scripts at about twenty cents each with better visual cues and cleaner handling of numbers, and can live with the weaker hooks.

Metric by metric

Blue bars: Kimi K3. Orange bars: Grok 4.6. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
88.3
87.9
Writing Craft & Clarity13% weight
88.5
86.5
Substance, Accuracy & Value15% weight
87.6
87.5
Continuity & Emotion14% weight
86.8
85.8
YouTube Best Practices12% weight
87.3
86.5
Hook Strength10% weight
90.2
80.8
Length Adherence8% weight
90.1
88.0
Slop Score (EQ-Bench + ours)5% weight
90.9
91.2
Visual Cue Quality4% weight
84.6
86.6

Everything else that differs

Kimi K3Grok 4.6
Overall / 10088.286.6
Writing Elo21802035
Run-to-run spread (± overall std)1.6802.530
Cost per script (USD)0.2600.204
Avg latency (s)234.3243.1
Open weightsYesNo

Full scorecards: Kimi K3 · Grok 4.6. How scoring works: methodology.

← All comparisons