Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.4 to 84.1.
Claude Opus 5 (max effort) wins this one outright. It sits first on the board at 2358.5 Elo; DeepSeek V4 Pro 0813 (max) ranks #39 at 1847.7, a gap of 510.8 points, and the confidence intervals do not overlap. The overall scores tell the same story: 90.37 against 84.09, a 6.28-point spread, and Opus is far steadier from script to script, with a standard deviation of 1.53 versus 6.49. The sharpest per-metric split is cue quality, where Opus leads 88.24 to 79.38. Continuity and emotion show the same pattern at 90.17 against 81.85. DeepSeek answers with price. It costs about two cents per script against roughly thirteen cents for Opus, nearly 6x cheaper, and it ships open weights you can host yourself. This is a quality-first matchup: Opus is the better writer everywhere, DeepSeek is the budget play with real but uneven output.
Pick Claude Opus 5 (max effort) if you want the board's best script quality, top-tier cue work, and output steady enough to publish without a rewrite pass.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights and strong hooks at about two cents per script and can tolerate run-to-run inconsistency.
Blue bars: Claude Opus 5 (max). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | DeepSeek V4 Pro 0813 (max) | |
|---|---|---|
| Overall / 100 | 90.4 | 84.1 |
| Writing Elo | 2358 | 1848 |
| Run-to-run spread (± overall std) | 1.530 | 6.490 |
| Cost per script (USD) | 0.127 | 0.022 |
| Avg latency (s) | 244.2 | 227.3 |
| Open weights | No | Yes |
Full scorecards: Claude Opus 5 (max) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.
← All comparisons