Towards AITowards AIToneBench

DeepSeek V4 Flash 0731 vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Flash 0731 leads overall, 84.2 to 83.4. Their confidence intervals overlap, so treat the order as close rather than settled.

DeepSeek V4 Flash 0731
#34 Elo 1955 · 84.2/100
Qwen3.8 Max
#38 Elo 1866 · 83.4/100
Cost / script
$0.005 vs $0.150
Human baseline
92.7 DeepSeek V4 Flash 0731 falls below it · Qwen3.8 Max falls below it

The verdict

Qwen3.8 Max at #38 and DeepSeek V4 Flash 0731 at #34. Adjacent on the board, 83.44 against 84.23, well under half a point. Their Elo intervals overlap, so I would treat the order as unsettled. They are not the same model though. Qwen wins length adherence enormously, 92.90 against 83.57, and anti-slop 87.78 against 79.86. DeepSeek wins voice 83.42 against 81.93, hooks 87.55 against 83.50, and continuity 80.43 against 78.08. So Qwen is the tidier, better-sized, cleaner draft; DeepSeek is the one that sounds more like a person. Then the part that actually decides it: DeepSeek costs half a cent a script and is open weights, Qwen costs 15 cents and is closed. Thirty times the price for the same score. Unless you specifically need Qwen's length control, this is an easy call.

Pick Qwen3.8 Max if exact length and low slop are what you need, and the price does not matter.
Pick DeepSeek V4 Flash 0731 for effectively the same score at a thirtieth of the cost, open weights, with better voice and hooks.

Metric by metric

Blue bars: DeepSeek V4 Flash 0731. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
85.4
81.9
Writing Craft & Clarity13% weight
85.3
83.4
Substance, Accuracy & Value15% weight
83.5
84.4
Continuity & Emotion14% weight
82.0
78.1
YouTube Best Practices12% weight
81.1
80.4
Hook Strength10% weight
87.5
83.5
Length Adherence8% weight
83.6
92.9
Slop Score (EQ-Bench + ours)5% weight
88.8
92.0
Visual Cue Quality4% weight
82.4
85.2

Everything else that differs

DeepSeek V4 Flash 0731Qwen3.8 Max
Overall / 10084.283.4
Writing Elo19551866
Run-to-run spread (± overall std)9.2405.620
Cost per script (USD)0.0050.150
Avg latency (s)289.0397.4
Open weightsYesNo

Full scorecards: DeepSeek V4 Flash 0731 · Qwen3.8 Max. How scoring works: methodology.

← All comparisons