Towards AITowards AIToneBench

Claude Opus 5 (max) vs Grok 4.5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 85.8.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
Grok 4.5
#30 Elo 2042 · 85.8/100
Cost / script
$0.120 vs $0.038
Human baseline
92.7 Claude Opus 5 (max) falls below it · Grok 4.5 falls below it

The verdict

Grok 4.5 is the cheap-and-fast option that stays respectable: #30, 85.8 overall, 42 seconds a script, and under four cents. Against Opus 5 at max effort that's 2041.9 Elo to 2649.3 and a 5.7 point quality gap, so nobody should pretend this is close. But the shape of the gap is worth reading. Grok holds voice and hooks reasonably, 87.07 and 87.96, and its substance score of 87.42 is solid. Where it falls apart is length adherence: 78.67 against 89.5. It does not hit a target word count, and on a long-form brief that is not a rounding error, that's a rewrite. Continuity is the other soft spot, 83.50 against 90.5. So the honest read is that Grok drafts fast and cheap and gets the ideas roughly right, then hands you something structurally off that you'll spend real time fixing. At a third of the cost and a sixth of the wall clock, that trade is fine for volume drafting and wrong for anything you're shipping.

Pick Opus 5 (max) if the draft needs to be close to final, especially on length and flow.
Pick Grok 4.5 if you're generating a lot of rough material fast and cheap, and you have an editor who'll fix the structure anyway.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
87.1
Writing Craft & Clarity13% weight
91.1
86.6
Substance, Accuracy & Value15% weight
90.6
87.4
Continuity & Emotion14% weight
90.4
83.5
YouTube Best Practices12% weight
89.6
84.4
Hook Strength10% weight
92.0
88.0
Length Adherence8% weight
89.7
78.7
Slop Score (EQ-Bench + ours)5% weight
93.3
90.2
Visual Cue Quality4% weight
89.6
86.5

Everything else that differs

Claude Opus 5 (max)Grok 4.5
Overall / 10090.885.8
Writing Elo26492042
Run-to-run spread (± overall std)1.2902.810
Cost per script (USD)0.1200.038
Avg latency (s)243.841.7
Open weightsNoNo

Full scorecards: Claude Opus 5 (max) · Grok 4.5. How scoring works: methodology.

← All comparisons