Towards AITowards AIToneBench

Kimi K3 (thinking) vs Grok 4.5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 85.8.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
Grok 4.5
#30 Elo 2042 · 85.8/100
Cost / script
$0.187 vs $0.038
Human baseline
92.7 Kimi K3 (thinking) falls below it · Grok 4.5 falls below it

The verdict

This one isn't close. Kimi K3 wins at 89.3 overall against Grok 4.5's 85.8, and the Elo intervals don't even touch, so no hedging needed. Credit where it's due first: Grok is fast, under 30 seconds per script, and at roughly 4 cents it costs about a fifth of Kimi. That combination is genuinely useful when you're iterating. But as a writer? Grok's biggest problem is discipline. It scores 78.7 on length adherence versus Kimi's 91.1, which in practice means scripts that overshoot or undershoot the target and need trimming by hand. Continuity and emotion tell the same story, 83.5 against 88.3: sections that sit next to each other instead of flowing into each other. Kimi leads on nearly every metric while staying open weights, which I appreciate, and its hooks at 90.2 are the best in this pair. If the script is going anywhere near publish, Kimi. If you're testing ten hook ideas in a minute, Grok earns its spot.

Pick Kimi K3 if the script is meant to publish, since it wins nearly every metric and stays open weights.
Pick Grok 4.5 if you're iterating fast on cheap drafts and will fix length and flow yourself.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
87.1
Writing Craft & Clarity13% weight
89.3
86.6
Substance, Accuracy & Value15% weight
89.2
87.4
Continuity & Emotion14% weight
88.3
83.5
YouTube Best Practices12% weight
87.8
84.4
Hook Strength10% weight
90.2
88.0
Length Adherence8% weight
91.1
78.7
Slop Score (EQ-Bench + ours)5% weight
91.6
90.2
Visual Cue Quality4% weight
87.7
86.5

Everything else that differs

Kimi K3 (thinking)Grok 4.5
Overall / 10089.285.8
Writing Elo24162042
Run-to-run spread (± overall std)1.6502.810
Cost per script (USD)0.1870.038
Avg latency (s)267.241.7
Open weightsYesNo

Full scorecards: Kimi K3 (thinking) · Grok 4.5. How scoring works: methodology.

← All comparisons