Towards AITowards AIToneBench

Kimi K3 vs Grok 4.5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 85.8.

Kimi K3
#8 Elo 2438 · 89.4/100
Grok 4.5
#30 Elo 2042 · 85.8/100
Cost / script
$0.263 vs $0.038
Human baseline
92.7 Kimi K3 falls below it · Grok 4.5 falls below it

The verdict

Kimi K3 at #8 against Grok 4.5 at #30, 89.39 overall to 85.8. Four and a quarter points, intervals well clear. The decisive metric is length adherence: Kimi 90.46, Grok 78.67. Fourteen points. Grok does not hit a target word count, and on long-form that alone justifies the ranking gap. Continuity is the other big one, 88.28 to 83.50. Where Grok stays competitive is substance at 87.42 against 89.6, and hooks at 87.96 against 89.36, so the thinking and the openings are not the problem. It's the shape of the finished piece. Now the other direction: Grok costs under four cents a script and returns in 42 seconds. Kimi costs 26 cents and takes over six minutes. That's seven times the price and ten times the wait for four points. If you're drafting at volume and editing anyway, that math favours Grok more than the leaderboard position suggests.

Pick Kimi K3 if the draft needs to hold its length and flow, and you want open weights.
Pick Grok 4.5 if throughput is the constraint. Seven times cheaper, ten times faster, and the ideas mostly land.

Metric by metric

Blue bars: Kimi K3. Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.6
87.1
Writing Craft & Clarity13% weight
89.3
86.6
Substance, Accuracy & Value15% weight
89.5
87.4
Continuity & Emotion14% weight
88.3
83.5
YouTube Best Practices12% weight
88.4
84.4
Hook Strength10% weight
89.9
88.0
Length Adherence8% weight
90.5
78.7
Slop Score (EQ-Bench + ours)5% weight
92.0
90.2
Visual Cue Quality4% weight
88.3
86.5

Everything else that differs

Kimi K3Grok 4.5
Overall / 10089.485.8
Writing Elo24382042
Run-to-run spread (± overall std)1.5702.810
Cost per script (USD)0.2630.038
Avg latency (s)367.141.7
Open weightsYesNo

Full scorecards: Kimi K3 · Grok 4.5. How scoring works: methodology.

← All comparisons