Towards AITowards AIToneBench

Claude Opus 5 (max) vs DeepSeek V4 Pro (xhigh)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 80.5.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
DeepSeek V4 Pro (xhigh)
#60 Elo 1601 · 80.5/100
Cost / script
$0.120 vs $0.014
Human baseline
92.7 Claude Opus 5 (max) falls below it · DeepSeek V4 Pro (xhigh) falls below it

The verdict

DeepSeek V4 Pro at xhigh is the cheapest model on this board by a wide margin, under a cent and a half per script, with open weights. It's also #60, and against Opus 5 that's 80.50 overall to 90.80. The single number that should decide this for you is anti-slop: 65.96 against Opus 5's 90.26. Twenty three points. That is not a polish gap, that's scripts full of AI-flavoured filler and unreliable numbers, and on a technical script wrong numbers are the expensive kind of wrong. Cue quality is the other collapse at 73.61, sixteen points down. What it does well is length, 80.47, and hooks at 84.61. So you get a well-sized draft with a decent opening that needs heavy line editing throughout. At eight times cheaper than Opus that can still pencil out for bulk drafting, as long as you know what you're signing up for on the edit side.

Pick Opus 5 (max) for anything with numbers in it, or anything going out without a careful line edit.
Pick DeepSeek V4 Pro (xhigh) if you need open weights at the lowest cost on the board and you have the editorial time to clean up the slop.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
82.2
Writing Craft & Clarity13% weight
91.1
82.2
Substance, Accuracy & Value15% weight
90.6
81.0
Continuity & Emotion14% weight
90.4
78.5
YouTube Best Practices12% weight
89.6
76.9
Hook Strength10% weight
92.0
84.6
Length Adherence8% weight
89.7
80.5
Slop Score (EQ-Bench + ours)5% weight
93.3
79.4
Visual Cue Quality4% weight
89.6
73.6

Everything else that differs

Claude Opus 5 (max)DeepSeek V4 Pro (xhigh)
Overall / 10090.880.5
Writing Elo26491601
Run-to-run spread (± overall std)1.2903.670
Cost per script (USD)0.1200.014
Avg latency (s)243.8163.2
Open weightsNoYes

Full scorecards: Claude Opus 5 (max) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.

← All comparisons