Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.4 to 88.1.
No contest on the board. Opus 5 at max effort sits first at 2358.5 Elo; Grok 4.6 high is seventeenth at 2144.5, and the confidence intervals, 2312 to 2409 against 2103 to 2181, are not close to touching. The overall scores tell the same story: 90.37 against 88.14, a 2.23-point gap, which is large for this leaderboard. The separation is widest on continuity and emotion, 90.17 to 86.27, and Opus also leads on cue quality, 88.24 to 86.43. Grok claims exactly one metric: length adherence, 88.26 to 87.2. The usual escape hatch, price, is shut here. Opus costs about thirteen cents per script (estimated) and Grok about twelve cents (exact), so Grok is not meaningfully cheaper. Neither is open weights. A rare case where the better writer costs about a penny more, and we would pay it.
Pick Claude Opus 5 (max effort) if you want the strongest writer on the board, especially for continuity and cue quality, and a roughly one-penny premium per script is irrelevant.
Pick Grok 4.6 (high) if you value exact, predictable billing and slightly tighter length adherence, and can accept a clear step down in overall writing quality.
Blue bars: Claude Opus 5 (max). Orange bars: Grok 4.6 (high). Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | Grok 4.6 (high) | |
|---|---|---|
| Overall / 100 | 90.4 | 88.1 |
| Writing Elo | 2358 | 2144 |
| Run-to-run spread (± overall std) | 1.530 | 1.430 |
| Cost per script (USD) | 0.127 | 0.115 |
| Avg latency (s) | 244.2 | 199.2 |
| Open weights | No | No |
Full scorecards: Claude Opus 5 (max) · Grok 4.6 (high). How scoring works: methodology.
← All comparisons