Towards AITowards AIToneBench

DeepSeek V4.1 Flash (max) vs GLM-5.3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 86.8 to 86.8. Their confidence intervals overlap, so treat the order as close rather than settled.

DeepSeek V4.1 Flash (max)
#22 Elo 2085 · 86.8/100
GLM-5.3
#27 Elo 2030 · 86.8/100
Cost / script
$0.012 vs $0.109
Human baseline
90.2 DeepSeek V4.1 Flash (max) falls below it · GLM-5.3 falls below it

The verdict

DeepSeek V4.1 Flash (max) has the stronger current result: #22 at 2085.1 Elo and 86.75 overall, compared with GLM-5.3 at #27, 2030.5 Elo, and 86.83 overall. Their 95% Elo confidence intervals overlap, so the exact ordering should be treated as uncertain (2037.20–2122.70 and 1982.10–2074.80). DeepSeek V4.1 Flash (max) has its clearest metric edges in Length Adherence (89.95 versus 87.50) and Visual Cue Quality (82.35 versus 81.43). GLM-5.3 counters on Substance, Accuracy & Value (85.93 versus 85.23) and Hook Strength (89.54 versus 88.88). At the measured run mix, DeepSeek V4.1 Flash (max) costs $0.012 per article versus $0.109 for GLM-5.3; DeepSeek V4.1 Flash (max) is the cheaper route. Both models publish open weights. On the current automated evidence, DeepSeek V4.1 Flash (max) is the stronger default; GLM-5.3 remains a defensible choice when its specific strengths, price, or deployment profile matter more than the headline rank.

Pick DeepSeek V4.1 Flash (max) when you prioritize the stronger current board result, length adherence, visual cue quality, open weights, and lower measured cost.
Pick GLM-5.3 when you prioritize substance, accuracy & value, hook strength, and open weights.

Metric by metric

Blue bars: DeepSeek V4.1 Flash (max). Orange bars: GLM-5.3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.4
87.9
Writing Craft & Clarity13% weight
87.3
87.2
Substance, Accuracy & Value15% weight
85.2
85.9
Continuity & Emotion14% weight
84.9
85.1
YouTube Best Practices12% weight
84.9
85.0
Hook Strength10% weight
88.9
89.5
Length Adherence8% weight
90.0
87.5
Slop Score (EQ-Bench + ours)5% weight
91.0
91.6
Visual Cue Quality4% weight
82.3
81.4

Everything else that differs

DeepSeek V4.1 Flash (max)GLM-5.3
Overall / 10086.886.8
Writing Elo20852030
Run-to-run spread (± overall std)6.9406.600
Cost per script (USD)0.0120.109
Avg latency (s)74.6301.9
Open weightsYesYes

Full scorecards: DeepSeek V4.1 Flash (max) · GLM-5.3. How scoring works: methodology.

← All comparisons