Towards AITowards AIToneBench

GLM-5.3 vs DeepSeek V4 Pro 0813 (max)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 86.8 to 84.0.

GLM-5.3
#24 Elo 2042 · 86.8/100
DeepSeek V4 Pro 0813 (max)
#37 Elo 1881 · 84.0/100
Cost / script
$0.109 vs $0.022
Human baseline
90.2 GLM-5.3 falls below it · DeepSeek V4 Pro 0813 (max) falls below it

The verdict

GLM-5.3 wins this one, and the intervals say it's real: GLM's Elo floor of 1989.0 sits above DeepSeek's ceiling of 1946.4. GLM ranks #24 at 2041.5 Elo against DeepSeek V4 Pro's #37 at 1880.7, a 160.8-point gap, and takes the overall score 86.83 to 84.04. It's a sweep across every metric, and the biggest gaps land where a script stops sounding like a script: numbers-and-slop at 86.75 vs 82.09, continuity and emotion at 85.11 vs 80.77, and YouTube best practices at 85.02 vs 81.13. The records agree, 866 wins and 8 losses for GLM against DeepSeek's 674 and 68. Both are open weights, so the tiebreaker is the bill. DeepSeek runs about two cents per script against GLM's roughly eleven, about 5x cheaper, and it's faster too. A 2.79-point overall gap at 5x less cost is a live trade if the draft gets rewritten anyway. If it ships close to as-written, GLM earns the premium.

Pick GLM-5.3 if you want the clear board winner, with stronger continuity, hooks, and fewer slop tells, at roughly eleven cents a script.
Pick DeepSeek V4 Pro 0813 (max) if you want drafts at about two cents each, delivered faster, and can accept a 2.79-point overall gap on scripts you plan to edit anyway.

Metric by metric

Blue bars: GLM-5.3. Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.9
84.7
Writing Craft & Clarity13% weight
87.2
84.9
Substance, Accuracy & Value15% weight
85.9
84.7
Continuity & Emotion14% weight
85.1
80.8
YouTube Best Practices12% weight
85.0
81.1
Hook Strength10% weight
89.5
86.8
Length Adherence8% weight
87.5
86.2
Slop Score (EQ-Bench + ours)5% weight
91.6
88.0
Visual Cue Quality4% weight
81.4
79.5

Everything else that differs

GLM-5.3DeepSeek V4 Pro 0813 (max)
Overall / 10086.884.0
Writing Elo20421881
Run-to-run spread (± overall std)6.6004.370
Cost per script (USD)0.1090.022
Avg latency (s)301.9195.2
Open weightsYesYes

Full scorecards: GLM-5.3 · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.

← All comparisons