Towards AITowards AIToneBench

Qwen3.7 Max (high) vs DeepSeek V4 Pro (xhigh)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.7 Max (high) leads overall, 81.5 to 80.5. Their confidence intervals overlap, so treat the order as close rather than settled.

Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
DeepSeek V4 Pro (xhigh)
#60 Elo 1601 · 80.5/100
Cost / script
$0.054 vs $0.014
Human baseline
92.7 Qwen3.7 Max (high) falls below it · DeepSeek V4 Pro (xhigh) falls below it

The verdict

On paper Qwen3.7 Max wins, but this is the one matchup where I would probably reach for the lower-ranked model. The Elo intervals overlap, the overall scores land just under a point apart, and the two split the metrics: Qwen is better behaved with numbers in spoken lines, 77.78 against DeepSeek's 65.96, while DeepSeek actually edges Qwen on tone and voice, 82.21 versus 80.91. That last one surprised me, and tone carries the heaviest weight in how I score these. Then the practical stuff tilts the same way. Qwen costs about three times more per script, four and a half cents against a cent and a half, and it is closed while DeepSeek ships open weights I can run where I want. When quality is this close, I let cost and openness break the tie. Qwen's cleaner stat handling is real and does save editing time. I just would not pay triple for what is close to a coin flip.

Pick Qwen3.7 Max (high) if cleaner number handling and length discipline save you real editing time at your volume.
Pick DeepSeek V4 Pro (xhigh) if you want roughly the same quality with a slightly better tone match, open weights, and a third of the price.

Metric by metric

Blue bars: Qwen3.7 Max (high). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
80.9
82.2
Writing Craft & Clarity13% weight
81.5
82.2
Substance, Accuracy & Value15% weight
81.2
81.0
Continuity & Emotion14% weight
79.2
78.5
YouTube Best Practices12% weight
80.7
76.9
Hook Strength10% weight
85.0
84.6
Length Adherence8% weight
83.3
80.5
Slop Score (EQ-Bench + ours)5% weight
85.9
79.4
Visual Cue Quality4% weight
78.4
73.6

Everything else that differs

Qwen3.7 Max (high)DeepSeek V4 Pro (xhigh)
Overall / 10081.580.5
Writing Elo16741601
Run-to-run spread (± overall std)2.5803.670
Cost per script (USD)0.0540.014
Avg latency (s)157.2163.2
Open weightsNoYes

Full scorecards: Qwen3.7 Max (high) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.

← All comparisons