Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.7 Max (default) leads overall, 81.7 to 80.5. Their confidence intervals overlap, so treat the order as close rather than settled.
Qwen3.7 Max at #50 and DeepSeek V4 Pro (xhigh) at #60, 81.73 against 80.50. One point apart on overall, which undersells how differently they fail. DeepSeek's anti-slop score is 65.96. That's the lowest number in any of these comparisons, eleven points below Qwen and twenty three below the top of the board. Filler and unreliable figures, throughout. Cue quality is the other collapse at 73.61 against Qwen's 80.24. What keeps DeepSeek's overall respectable is voice at 82.21 and length at 80.47, both slightly ahead of Qwen. So it sounds fine and comes out the right length while getting details wrong, which is arguably the more dangerous failure mode for a technical script. DeepSeek is open weights and under a cent and a half, roughly a quarter of Qwen's price. That's a real argument for bulk drafting. It is not an argument for anything with numbers in it.
Pick Qwen3.7 Max if the script has figures or visual direction that need to be right.
Pick DeepSeek V4 Pro (xhigh) if you want open weights at the lowest cost on the board and you're line editing everything anyway.
Blue bars: Qwen3.7 Max (default). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.
| Qwen3.7 Max (default) | DeepSeek V4 Pro (xhigh) | |
|---|---|---|
| Overall / 100 | 81.7 | 80.5 |
| Writing Elo | 1699 | 1601 |
| Run-to-run spread (± overall std) | 3.060 | 3.670 |
| Cost per script (USD) | 0.054 | 0.014 |
| Avg latency (s) | 156.8 | 163.2 |
| Open weights | No | Yes |
Full scorecards: Qwen3.7 Max (default) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.
← All comparisons