Towards AITowards AIToneBench

Claude Opus 5 (max) vs DeepSeek V4 Flash 0731

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 84.2.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
DeepSeek V4 Flash 0731
#34 Elo 1955 · 84.2/100
Cost / script
$0.120 vs $0.005
Human baseline
92.7 Claude Opus 5 (max) falls below it · DeepSeek V4 Flash 0731 falls below it

The verdict

The 0731 build is the biggest single-model improvement I have measured on this board. DeepSeek V4 Flash went from #83 and 76.59 overall to #34 and 84.23. Seven and a half points, forty nine places, and the new build is cheaper than the old one. Against Opus 5 at max effort it is still six and a half points back, 90.80 to 84.23. Its hooks are genuinely good at 87.55, about four behind the board leader, and its voice at 85.44 is better than its rank suggests. What holds it down is structure: YouTube best practices 81.06 against 89.56, continuity 82.02 against 90.44, both eight to nine points. It opens well and loses the thread. Now the number that matters most: half a cent per script against Opus 5's 12 cents. About twenty five times cheaper, open weights, for six and a half points. For bulk drafting where a human owns the outline, that math is very hard to argue with.

Pick Opus 5 (max) when the model has to carry the whole script and the draft needs to be near final.
Pick DeepSeek V4 Flash 0731 for volume. Open weights, half a cent a script, strong hooks, and you fix the structure yourself.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
85.4
Writing Craft & Clarity13% weight
91.1
85.3
Substance, Accuracy & Value15% weight
90.6
83.5
Continuity & Emotion14% weight
90.4
82.0
YouTube Best Practices12% weight
89.6
81.1
Hook Strength10% weight
92.0
87.5
Length Adherence8% weight
89.7
83.6
Slop Score (EQ-Bench + ours)5% weight
93.3
88.8
Visual Cue Quality4% weight
89.6
82.4

Everything else that differs

Claude Opus 5 (max)DeepSeek V4 Flash 0731
Overall / 10090.884.2
Writing Elo26491955
Run-to-run spread (± overall std)1.2909.240
Cost per script (USD)0.1200.005
Avg latency (s)243.8289.0
Open weightsNoYes

Full scorecards: Claude Opus 5 (max) · DeepSeek V4 Flash 0731. How scoring works: methodology.

← All comparisons