Towards AITowards AIToneBench

Grok 4.5 vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 82.4.

Grok 4.5
#30 Elo 2042 · 85.8/100
GLM-5
#43 Elo 1783 · 82.4/100
Cost / script
$0.038 vs $0.023
Human baseline
92.7 Grok 4.5 falls below it · GLM-5 falls below it

The verdict

A closer fight than the ranks suggest, and the first thing to know is that cost disappears from this decision: both land around two cents per script, GLM's estimated. Grok 4.5 wins, and even with only two points between them overall, the Elo intervals don't overlap, so I'll trust the order. Where they actually split is numbers. Grok scores 85.6 on our anti-slop and numbers check while GLM-5 sits at 77.3, meaning GLM is the one reciting stat sheets at you. GLM punches back on length adherence, 87.1 to Grok's 78.7, and Grok's variance there is high, so expect the occasional draft that overshoots badly. On tone the gap is a couple of points with standard deviations to match, so honestly, inside the noise. Two practical tiebreakers: Grok returns a script in under 30 seconds, the fastest here, while GLM is open weights, which Grok can't offer. Both trail my reference scripts by a wide margin, to be clear. For quality per dollar, Grok, narrowly.

Pick Grok 4.5 if you want cleaner number handling and near-instant drafts at the same two-cent price.
Pick GLM-5 if open weights matter and you'd rather fix the occasional stat dump than fix length overruns.

Metric by metric

Blue bars: Grok 4.5. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.1
83.1
Writing Craft & Clarity13% weight
86.6
84.0
Substance, Accuracy & Value15% weight
87.4
84.0
Continuity & Emotion14% weight
83.5
79.7
YouTube Best Practices12% weight
84.4
78.3
Hook Strength10% weight
88.0
86.5
Length Adherence8% weight
78.7
78.6
Slop Score (EQ-Bench + ours)5% weight
90.2
85.3
Visual Cue Quality4% weight
86.5
83.8

Everything else that differs

Grok 4.5GLM-5
Overall / 10085.882.4
Writing Elo20421783
Run-to-run spread (± overall std)2.8103.760
Cost per script (USD)0.0380.023
Avg latency (s)41.7117.3
Open weightsNoYes

Full scorecards: Grok 4.5 · GLM-5. How scoring works: methodology.

← All comparisons