Towards AITowards AIToneBench

Qwen3.8 Max vs DeepSeek V4 Flash 0731

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Flash 0731 leads overall, 84.2 to 83.4. Their confidence intervals overlap, so treat the order as close rather than settled.

Qwen3.8 Max
#38 Elo 1866 · 83.4/100
DeepSeek V4 Flash 0731
#34 Elo 1955 · 84.2/100
Cost / script
$0.150 vs $0.005
Human baseline
92.7 Qwen3.8 Max falls below it · DeepSeek V4 Flash 0731 falls below it

The verdict

Qwen3.8 Max at #38 and DeepSeek V4 Flash 0731 at #34. Adjacent on the board, 83.44 against 84.23, well under half a point. Their Elo intervals overlap, so I would treat the order as unsettled. They are not the same model though. Qwen wins length adherence enormously, 92.90 against 83.57, and anti-slop 87.78 against 79.86. DeepSeek wins voice 83.42 against 81.93, hooks 87.55 against 83.50, and continuity 80.43 against 78.08. So Qwen is the tidier, better-sized, cleaner draft; DeepSeek is the one that sounds more like a person. Then the part that actually decides it: DeepSeek costs half a cent a script and is open weights, Qwen costs 15 cents and is closed. Thirty times the price for the same score. Unless you specifically need Qwen's length control, this is an easy call.

Pick Qwen3.8 Max if exact length and low slop are what you need, and the price does not matter.
Pick DeepSeek V4 Flash 0731 for effectively the same score at a thirtieth of the cost, open weights, with better voice and hooks.

Metric by metric

Blue bars: Qwen3.8 Max. Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
81.9
85.4
Writing Craft & Clarity13% weight
83.4
85.3
Substance, Accuracy & Value15% weight
84.4
83.5
Continuity & Emotion14% weight
78.1
82.0
YouTube Best Practices12% weight
80.4
81.1
Hook Strength10% weight
83.5
87.5
Length Adherence8% weight
92.9
83.6
Slop Score (EQ-Bench + ours)5% weight
92.0
88.8
Visual Cue Quality4% weight
85.2
82.4

Everything else that differs

Qwen3.8 MaxDeepSeek V4 Flash 0731
Overall / 10083.484.2
Writing Elo18661955
Run-to-run spread (± overall std)5.6209.240
Cost per script (USD)0.1500.005
Avg latency (s)397.4289.0
Open weightsNoYes

Full scorecards: Qwen3.8 Max · DeepSeek V4 Flash 0731. How scoring works: methodology.

← All comparisons