Towards AITowards AIToneBench

DeepSeek V4.1 Flash (max) vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4.1 Flash (max) leads overall, 86.8 to 81.1.

DeepSeek V4.1 Flash (max)
#22 Elo 2085 · 86.8/100
Qwen3.8 Max
#52 Elo 1692 · 81.1/100
Cost / script
$0.012 vs $0.164
Human baseline
90.2 DeepSeek V4.1 Flash (max) falls below it · Qwen3.8 Max falls below it

The verdict

DeepSeek V4.1 Flash (max) has the stronger current result: #22 at 2085.1 Elo and 86.75 overall, compared with Qwen3.8 Max at #52, 1692.0 Elo, and 81.06 overall. Their 95% Elo confidence intervals do not overlap, which supports the directional ordering (2037.20–2122.70 and 1622.20–1764.00). DeepSeek V4.1 Flash (max) has its clearest metric edges in Continuity & Emotion (84.93 versus 76.22) and YouTube Best Practices (84.92 versus 77.76). Qwen3.8 Max does not lead an individual published metric in this pairing. At the measured run mix, DeepSeek V4.1 Flash (max) costs $0.012 per article versus $0.164 for Qwen3.8 Max; DeepSeek V4.1 Flash (max) is the cheaper route. DeepSeek V4.1 Flash (max) is the open-weights option; Qwen3.8 Max is closed. On the current automated evidence, DeepSeek V4.1 Flash (max) is the stronger default; Qwen3.8 Max remains a defensible choice when its specific strengths, price, or deployment profile matter more than the headline rank.

Pick DeepSeek V4.1 Flash (max) when you prioritize the stronger current board result, continuity & emotion, youtube best practices, open weights, and lower measured cost.
Pick Qwen3.8 Max when you prioritize its particular deployment route.

Metric by metric

Blue bars: DeepSeek V4.1 Flash (max). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.4
80.7
Writing Craft & Clarity13% weight
87.3
81.0
Substance, Accuracy & Value15% weight
85.2
81.3
Continuity & Emotion14% weight
84.9
76.2
YouTube Best Practices12% weight
84.9
77.8
Hook Strength10% weight
88.9
83.4
Length Adherence8% weight
90.0
86.3
Slop Score (EQ-Bench + ours)5% weight
91.0
90.4
Visual Cue Quality4% weight
82.3
80.6

Everything else that differs

DeepSeek V4.1 Flash (max)Qwen3.8 Max
Overall / 10086.881.1
Writing Elo20851692
Run-to-run spread (± overall std)6.9407.890
Cost per script (USD)0.0120.164
Avg latency (s)74.6429.0
Open weightsYesNo

Full scorecards: DeepSeek V4.1 Flash (max) · Qwen3.8 Max. How scoring works: methodology.

← All comparisons