Towards AITowards AIToneBench

GLM-5 vs DeepSeek V4 Pro (xhigh)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.4 to 80.5.

GLM-5
#43 Elo 1783 · 82.4/100
DeepSeek V4 Pro (xhigh)
#60 Elo 1601 · 80.5/100
Cost / script
$0.023 vs $0.014
Human baseline
92.7 GLM-5 falls below it · DeepSeek V4 Pro (xhigh) falls below it

The verdict

Two open models, both under two cents a script, neither anywhere near the top of my board. That makes it a genuinely fair fight. GLM-5 wins it, though I want to be honest about the margin. The average scores are close, 82.42 against 80.50, with enough spread that they nearly touch. But the rubric-derived Elo tells a cleaner story: GLM's confidence interval sits fully above DeepSeek's across the same scored task set. The gap I can point at is visual cues, 83.79 versus 73.61. GLM gives my editor usable [SHOW:] directions; DeepSeek's read like afterthoughts. Both share the same bad habit of leaving raw stats in the spoken lines, GLM just does it less. At this price neither is a final-draft model for me, they are drafting engines. If I am picking a drafting engine at a cent or two per script, GLM is the one I load first.

Pick GLM-5 if you want the stronger budget open-weights draft, especially for visual cues your editor can actually use.
Pick DeepSeek V4 Pro (xhigh) if you are already running it and a slightly cheaper, roughly comparable draft is enough.

Metric by metric

Blue bars: GLM-5. Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.1
82.2
Writing Craft & Clarity13% weight
84.0
82.2
Substance, Accuracy & Value15% weight
84.0
81.0
Continuity & Emotion14% weight
79.7
78.5
YouTube Best Practices12% weight
78.3
76.9
Hook Strength10% weight
86.5
84.6
Length Adherence8% weight
78.6
80.5
Slop Score (EQ-Bench + ours)5% weight
85.3
79.4
Visual Cue Quality4% weight
83.8
73.6

Everything else that differs

GLM-5DeepSeek V4 Pro (xhigh)
Overall / 10082.480.5
Writing Elo17831601
Run-to-run spread (± overall std)3.7603.670
Cost per script (USD)0.0230.014
Avg latency (s)117.3163.2
Open weightsYesYes

Full scorecards: GLM-5 · DeepSeek V4 Pro (xhigh). How scoring works: methodology.

← All comparisons