Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.7 Max (high) leads overall, 81.5 to 80.5. Their confidence intervals overlap, so treat the order as close rather than settled.
On paper Qwen3.7 Max wins, but this is the one matchup where I would probably reach for the lower-ranked model. The Elo intervals overlap, the overall scores land just under a point apart, and the two split the metrics: Qwen is better behaved with numbers in spoken lines, 77.78 against DeepSeek's 65.96, while DeepSeek actually edges Qwen on tone and voice, 82.21 versus 80.91. That last one surprised me, and tone carries the heaviest weight in how I score these. Then the practical stuff tilts the same way. Qwen costs about three times more per script, four and a half cents against a cent and a half, and it is closed while DeepSeek ships open weights I can run where I want. When quality is this close, I let cost and openness break the tie. Qwen's cleaner stat handling is real and does save editing time. I just would not pay triple for what is close to a coin flip.
Pick Qwen3.7 Max (high) if cleaner number handling and length discipline save you real editing time at your volume.
Pick DeepSeek V4 Pro (xhigh) if you want roughly the same quality with a slightly better tone match, open weights, and a third of the price.
Blue bars: Qwen3.7 Max (high). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.
| Qwen3.7 Max (high) | DeepSeek V4 Pro (xhigh) | |
|---|---|---|
| Overall / 100 | 81.5 | 80.5 |
| Writing Elo | 1674 | 1601 |
| Run-to-run spread (± overall std) | 2.580 | 3.670 |
| Cost per script (USD) | 0.054 | 0.014 |
| Avg latency (s) | 157.2 | 163.2 |
| Open weights | No | Yes |
Full scorecards: Qwen3.7 Max (high) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.
← All comparisons