Towards AITowards AIToneBench

Qwen3.7 Max (default) vs DeepSeek V4 Pro (xhigh)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.7 Max (default) leads overall, 81.7 to 80.5. Their confidence intervals overlap, so treat the order as close rather than settled.

Qwen3.7 Max (default)
#50 Elo 1699 · 81.7/100
DeepSeek V4 Pro (xhigh)
#60 Elo 1601 · 80.5/100
Cost / script
$0.054 vs $0.014
Human baseline
92.7 Qwen3.7 Max (default) falls below it · DeepSeek V4 Pro (xhigh) falls below it

The verdict

Qwen3.7 Max at #50 and DeepSeek V4 Pro (xhigh) at #60, 81.73 against 80.50. One point apart on overall, which undersells how differently they fail. DeepSeek's anti-slop score is 65.96. That's the lowest number in any of these comparisons, eleven points below Qwen and twenty three below the top of the board. Filler and unreliable figures, throughout. Cue quality is the other collapse at 73.61 against Qwen's 80.24. What keeps DeepSeek's overall respectable is voice at 82.21 and length at 80.47, both slightly ahead of Qwen. So it sounds fine and comes out the right length while getting details wrong, which is arguably the more dangerous failure mode for a technical script. DeepSeek is open weights and under a cent and a half, roughly a quarter of Qwen's price. That's a real argument for bulk drafting. It is not an argument for anything with numbers in it.

Pick Qwen3.7 Max if the script has figures or visual direction that need to be right.
Pick DeepSeek V4 Pro (xhigh) if you want open weights at the lowest cost on the board and you're line editing everything anyway.

Metric by metric

Blue bars: Qwen3.7 Max (default). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
81.2
82.2
Writing Craft & Clarity13% weight
81.7
82.2
Substance, Accuracy & Value15% weight
80.9
81.0
Continuity & Emotion14% weight
79.0
78.5
YouTube Best Practices12% weight
80.8
76.9
Hook Strength10% weight
84.2
84.6
Length Adherence8% weight
85.8
80.5
Slop Score (EQ-Bench + ours)5% weight
86.0
79.4
Visual Cue Quality4% weight
80.2
73.6

Everything else that differs

Qwen3.7 Max (default)DeepSeek V4 Pro (xhigh)
Overall / 10081.780.5
Writing Elo16991601
Run-to-run spread (± overall std)3.0603.670
Cost per script (USD)0.0540.014
Avg latency (s)156.8163.2
Open weightsNoYes

Full scorecards: Qwen3.7 Max (default) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.

← All comparisons