Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Pro 0813 (max) leads overall, 84.1 to 83.3. Their confidence intervals overlap, so treat the order as close rather than settled.
DeepSeek V4 Pro 0813 (max) takes this one, sitting at #39 with an elo of 1847.7 against Qwen3.8 Max at #42 and 1798.7. The 49.0-point elo gap looks decisive, but the confidence intervals overlap, so treat the ordering as provisional. The overall scores are nearly level: 84.09 to 83.35, a 0.74-point difference. The profiles diverge more than the totals suggest. DeepSeek writes the stronger openings, with hook strength at 87.99 against 83.95, and holds the better register at 85.57 on tone versus 81.81. Qwen3.8 Max is the more disciplined drafter: 90.63 on length adherence to DeepSeek's 81.21, and 84.78 on cue quality to 79.38. Cost separates them cleanly. DeepSeek runs about two cents per script; Qwen3.8 Max costs about fifteen cents, close to 7x more. DeepSeek is also open weights, which Qwen3.8 Max is not.
Pick DeepSeek V4 Pro 0813 (max) if you want stronger hooks and tone at a fraction of the price, plus open weights you can host yourself.
Pick Qwen3.8 Max if hitting target length and clean structural cues matter more to you than cost, since it leads clearly on both.
Blue bars: DeepSeek V4 Pro 0813 (max). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| DeepSeek V4 Pro 0813 (max) | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 84.1 | 83.3 |
| Writing Elo | 1848 | 1799 |
| Run-to-run spread (± overall std) | 6.490 | 5.430 |
| Cost per script (USD) | 0.022 | 0.150 |
| Avg latency (s) | 227.3 | 395.6 |
| Open weights | Yes | No |
Full scorecards: DeepSeek V4 Pro 0813 (max) · Qwen3.8 Max. How scoring works: methodology.
← All comparisons