Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 84.2.
The 0731 build is the biggest single-model improvement I have measured on this board. DeepSeek V4 Flash went from #83 and 76.59 overall to #34 and 84.23. Seven and a half points, forty nine places, and the new build is cheaper than the old one. Against Opus 5 at max effort it is still six and a half points back, 90.80 to 84.23. Its hooks are genuinely good at 87.55, about four behind the board leader, and its voice at 85.44 is better than its rank suggests. What holds it down is structure: YouTube best practices 81.06 against 89.56, continuity 82.02 against 90.44, both eight to nine points. It opens well and loses the thread. Now the number that matters most: half a cent per script against Opus 5's 12 cents. About twenty five times cheaper, open weights, for six and a half points. For bulk drafting where a human owns the outline, that math is very hard to argue with.
Pick Opus 5 (max) when the model has to carry the whole script and the draft needs to be near final.
Pick DeepSeek V4 Flash 0731 for volume. Open weights, half a cent a script, strong hooks, and you fix the structure yourself.
Blue bars: Claude Opus 5 (max). Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | DeepSeek V4 Flash 0731 | |
|---|---|---|
| Overall / 100 | 90.8 | 84.2 |
| Writing Elo | 2649 | 1955 |
| Run-to-run spread (± overall std) | 1.290 | 9.240 |
| Cost per script (USD) | 0.120 | 0.005 |
| Avg latency (s) | 243.8 | 289.0 |
| Open weights | No | Yes |
Full scorecards: Claude Opus 5 (max) · DeepSeek V4 Flash 0731. How scoring works: methodology.
← All comparisons