Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5 leads overall, 82.4 to 80.2.
An open model at around 2 cents a script beating Gemini 3.1 Pro on a writing benchmark would've sounded strange not long ago, but here we are. GLM-5 wins this cleanly: 82.4 vs 80.2 overall, and the Elo intervals don't overlap, so the gap is real. The wins are in the places I care about: voice match (83.1 vs 80.0) and length discipline, where Gemini drifts to 75.9 while GLM holds 78.6. Gemini fights back in two spots, and they're real. It handles numbers more cleanly (80.1 vs GLM's 76.5, honestly GLM's worst habit on this benchmark), and it returns drafts more than twice as fast. If you're iterating interactively, that speed matters. But GLM costs roughly 2 cents per script, estimated, against Gemini's 12, and you can run the weights yourself. Paying about seven times more for the lower-scoring script is a hard sell. Unless latency or an existing Google stack decides it for you, GLM takes this matchup.
Pick GLM-5 if you want the higher-scoring script for roughly 2 cents and the option to run the weights yourself.
Pick Gemini 3.1 Pro if turnaround speed and cleaner number handling matter more to you than the overall gap and the roughly seven-fold price difference.
Blue bars: GLM-5. Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.
| GLM-5 | Gemini 3.1 Pro (default) | |
|---|---|---|
| Overall / 100 | 82.4 | 80.2 |
| Writing Elo | 1783 | 1577 |
| Run-to-run spread (± overall std) | 3.760 | 2.420 |
| Cost per script (USD) | 0.023 | 0.118 |
| Avg latency (s) | 117.3 | 63.8 |
| Open weights | Yes | No |
Full scorecards: GLM-5 · Gemini 3.1 Pro (default). How scoring works: methodology.
← All comparisons