Towards AITowards AIToneBench

GLM-5 vs Gemini 3.1 Pro (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.4 to 80.2.

GLM-5
#43 Elo 1783 · 82.4/100
Gemini 3.1 Pro (default)
#63 Elo 1577 · 80.2/100
Cost / script
$0.023 vs $0.118
Human baseline
92.7 GLM-5 falls below it · Gemini 3.1 Pro (default) falls below it

The verdict

An open model at around 2 cents a script beating Gemini 3.1 Pro on a writing benchmark would've sounded strange not long ago, but here we are. GLM-5 wins this cleanly: 82.4 vs 80.2 overall, and the Elo intervals don't overlap, so the gap is real. The wins are in the places I care about: voice match (83.1 vs 80.0) and length discipline, where Gemini drifts to 75.9 while GLM holds 78.6. Gemini fights back in two spots, and they're real. It handles numbers more cleanly (80.1 vs GLM's 76.5, honestly GLM's worst habit on this benchmark), and it returns drafts more than twice as fast. If you're iterating interactively, that speed matters. But GLM costs roughly 2 cents per script, estimated, against Gemini's 12, and you can run the weights yourself. Paying about seven times more for the lower-scoring script is a hard sell. Unless latency or an existing Google stack decides it for you, GLM takes this matchup.

Pick GLM-5 if you want the higher-scoring script for roughly 2 cents and the option to run the weights yourself.
Pick Gemini 3.1 Pro if turnaround speed and cleaner number handling matter more to you than the overall gap and the roughly seven-fold price difference.

Metric by metric

Blue bars: GLM-5. Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.1
79.9
Writing Craft & Clarity13% weight
84.0
80.8
Substance, Accuracy & Value15% weight
84.0
80.8
Continuity & Emotion14% weight
79.7
78.4
YouTube Best Practices12% weight
78.3
78.9
Hook Strength10% weight
86.5
82.3
Length Adherence8% weight
78.6
75.9
Slop Score (EQ-Bench + ours)5% weight
85.3
87.7
Visual Cue Quality4% weight
83.8
81.7

Everything else that differs

GLM-5Gemini 3.1 Pro (default)
Overall / 10082.480.2
Writing Elo17831577
Run-to-run spread (± overall std)3.7602.420
Cost per script (USD)0.0230.118
Avg latency (s)117.363.8
Open weightsYesNo

Full scorecards: GLM-5 · Gemini 3.1 Pro (default). How scoring works: methodology.

← All comparisons