Towards AITowards AIToneBench

DeepSeek V4 Pro 0813 (max) vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Pro 0813 (max) leads overall, 84.1 to 81.9.

DeepSeek V4 Pro 0813 (max)
#38 Elo 1848 · 84.1/100
GLM-5
#49 Elo 1691 · 81.9/100
Cost / script
$0.022 vs $0.024
Human baseline
90.3 DeepSeek V4 Pro 0813 (max) falls below it · GLM-5 falls below it

The verdict

DeepSeek V4 Pro 0813 (max) takes this matchup cleanly. It sits at #39 with an elo of 1847.7 against GLM-5 at #55 and 1690.6, a gap of 157.1 points, and the confidence intervals do not overlap. The overall scores are much closer: 84.09 versus 81.91, a spread of 2.18, and GLM-5 is the steadier performer with the lower deviation of the two. The separation comes from the slop family. DeepSeek posts 83.59 on anti-slop numbers where GLM-5 drops to 75.17, and it leads on slop, 88.75 to 85.04. GLM-5 answers with better cue quality, 81.97 to 79.38, and a hair more substance accuracy. Cost is a wash: both land around two cents per script, with DeepSeek a shade cheaper. Both ship open weights, so licensing settles nothing. On this evidence, DeepSeek is the default pick unless cue discipline matters most to you.

Pick DeepSeek V4 Pro 0813 (max) if you want the clearly stronger head-to-head record and much cleaner numbers handling, with slop control and hook strength both a step above GLM-5 at essentially the same price.
Pick GLM-5 if you value tighter run-to-run consistency and better cue quality, and can accept a lower board position in exchange.

Metric by metric

Blue bars: DeepSeek V4 Pro 0813 (max). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
85.6
82.5
Writing Craft & Clarity13% weight
85.6
83.5
Substance, Accuracy & Value15% weight
83.4
83.5
Continuity & Emotion14% weight
81.8
78.8
YouTube Best Practices12% weight
81.9
78.1
Hook Strength10% weight
88.0
86.2
Length Adherence8% weight
81.2
78.7
Slop Score (EQ-Bench + ours)5% weight
88.8
85.0
Visual Cue Quality4% weight
79.4
82.0

Everything else that differs

DeepSeek V4 Pro 0813 (max)GLM-5
Overall / 10084.181.9
Writing Elo18481691
Run-to-run spread (± overall std)6.4903.590
Cost per script (USD)0.0220.024
Avg latency (s)227.3121.2
Open weightsYesYes

Full scorecards: DeepSeek V4 Pro 0813 (max) · GLM-5. How scoring works: methodology.

← All comparisons