Towards AITowards AIToneBench

Claude Opus 5 (max) vs DeepSeek V4 Pro 0813 (max)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.4 to 84.1.

Claude Opus 5 (max)
#1 Elo 2358 · 90.4/100
DeepSeek V4 Pro 0813 (max)
#38 Elo 1848 · 84.1/100
Cost / script
$0.127 vs $0.022
Human baseline
90.3 Claude Opus 5 (max) exceeds it · DeepSeek V4 Pro 0813 (max) falls below it

The verdict

Claude Opus 5 (max effort) wins this one outright. It sits first on the board at 2358.5 Elo; DeepSeek V4 Pro 0813 (max) ranks #39 at 1847.7, a gap of 510.8 points, and the confidence intervals do not overlap. The overall scores tell the same story: 90.37 against 84.09, a 6.28-point spread, and Opus is far steadier from script to script, with a standard deviation of 1.53 versus 6.49. The sharpest per-metric split is cue quality, where Opus leads 88.24 to 79.38. Continuity and emotion show the same pattern at 90.17 against 81.85. DeepSeek answers with price. It costs about two cents per script against roughly thirteen cents for Opus, nearly 6x cheaper, and it ships open weights you can host yourself. This is a quality-first matchup: Opus is the better writer everywhere, DeepSeek is the budget play with real but uneven output.

Pick Claude Opus 5 (max effort) if you want the board's best script quality, top-tier cue work, and output steady enough to publish without a rewrite pass.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights and strong hooks at about two cents per script and can tolerate run-to-run inconsistency.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
90.9
85.6
Writing Craft & Clarity13% weight
91.0
85.6
Substance, Accuracy & Value15% weight
90.3
83.4
Continuity & Emotion14% weight
90.2
81.8
YouTube Best Practices12% weight
89.6
81.9
Hook Strength10% weight
91.9
88.0
Length Adherence8% weight
87.2
81.2
Slop Score (EQ-Bench + ours)5% weight
93.0
88.8
Visual Cue Quality4% weight
88.2
79.4

Everything else that differs

Claude Opus 5 (max)DeepSeek V4 Pro 0813 (max)
Overall / 10090.484.1
Writing Elo23581848
Run-to-run spread (± overall std)1.5306.490
Cost per script (USD)0.1270.022
Avg latency (s)244.2227.3
Open weightsNoYes

Full scorecards: Claude Opus 5 (max) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.

← All comparisons