Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 80.5.
DeepSeek V4 Pro at xhigh is the cheapest model on this board by a wide margin, under a cent and a half per script, with open weights. It's also #60, and against Opus 5 that's 80.50 overall to 90.80. The single number that should decide this for you is anti-slop: 65.96 against Opus 5's 90.26. Twenty three points. That is not a polish gap, that's scripts full of AI-flavoured filler and unreliable numbers, and on a technical script wrong numbers are the expensive kind of wrong. Cue quality is the other collapse at 73.61, sixteen points down. What it does well is length, 80.47, and hooks at 84.61. So you get a well-sized draft with a decent opening that needs heavy line editing throughout. At eight times cheaper than Opus that can still pencil out for bulk drafting, as long as you know what you're signing up for on the edit side.
Pick Opus 5 (max) for anything with numbers in it, or anything going out without a careful line edit.
Pick DeepSeek V4 Pro (xhigh) if you need open weights at the lowest cost on the board and you have the editorial time to clean up the slop.
Blue bars: Claude Opus 5 (max). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | DeepSeek V4 Pro (xhigh) | |
|---|---|---|
| Overall / 100 | 90.8 | 80.5 |
| Writing Elo | 2649 | 1601 |
| Run-to-run spread (± overall std) | 1.290 | 3.670 |
| Cost per script (USD) | 0.120 | 0.014 |
| Avg latency (s) | 243.8 | 163.2 |
| Open weights | No | Yes |
Full scorecards: Claude Opus 5 (max) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.
← All comparisons