Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 85.8.
Grok 4.5 is the cheap-and-fast option that stays respectable: #30, 85.8 overall, 42 seconds a script, and under four cents. Against Opus 5 at max effort that's 2041.9 Elo to 2649.3 and a 5.7 point quality gap, so nobody should pretend this is close. But the shape of the gap is worth reading. Grok holds voice and hooks reasonably, 87.07 and 87.96, and its substance score of 87.42 is solid. Where it falls apart is length adherence: 78.67 against 89.5. It does not hit a target word count, and on a long-form brief that is not a rounding error, that's a rewrite. Continuity is the other soft spot, 83.50 against 90.5. So the honest read is that Grok drafts fast and cheap and gets the ideas roughly right, then hands you something structurally off that you'll spend real time fixing. At a third of the cost and a sixth of the wall clock, that trade is fine for volume drafting and wrong for anything you're shipping.
Pick Opus 5 (max) if the draft needs to be close to final, especially on length and flow.
Pick Grok 4.5 if you're generating a lot of rough material fast and cheap, and you have an editor who'll fix the structure anyway.
Blue bars: Claude Opus 5 (max). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | Grok 4.5 | |
|---|---|---|
| Overall / 100 | 90.8 | 85.8 |
| Writing Elo | 2649 | 2042 |
| Run-to-run spread (± overall std) | 1.290 | 2.810 |
| Cost per script (USD) | 0.120 | 0.038 |
| Avg latency (s) | 243.8 | 41.7 |
| Open weights | No | No |
Full scorecards: Claude Opus 5 (max) · Grok 4.5. How scoring works: methodology.
← All comparisons