Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Flash 0731 leads overall, 84.2 to 83.4. Their confidence intervals overlap, so treat the order as close rather than settled.
Qwen3.8 Max at #38 and DeepSeek V4 Flash 0731 at #34. Adjacent on the board, 83.44 against 84.23, well under half a point. Their Elo intervals overlap, so I would treat the order as unsettled. They are not the same model though. Qwen wins length adherence enormously, 92.90 against 83.57, and anti-slop 87.78 against 79.86. DeepSeek wins voice 83.42 against 81.93, hooks 87.55 against 83.50, and continuity 80.43 against 78.08. So Qwen is the tidier, better-sized, cleaner draft; DeepSeek is the one that sounds more like a person. Then the part that actually decides it: DeepSeek costs half a cent a script and is open weights, Qwen costs 15 cents and is closed. Thirty times the price for the same score. Unless you specifically need Qwen's length control, this is an easy call.
Pick Qwen3.8 Max if exact length and low slop are what you need, and the price does not matter.
Pick DeepSeek V4 Flash 0731 for effectively the same score at a thirtieth of the cost, open weights, with better voice and hooks.
Blue bars: DeepSeek V4 Flash 0731. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| DeepSeek V4 Flash 0731 | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 84.2 | 83.4 |
| Writing Elo | 1955 | 1866 |
| Run-to-run spread (± overall std) | 9.240 | 5.620 |
| Cost per script (USD) | 0.005 | 0.150 |
| Avg latency (s) | 289.0 | 397.4 |
| Open weights | Yes | No |
Full scorecards: DeepSeek V4 Flash 0731 · Qwen3.8 Max. How scoring works: methodology.
← All comparisons