Towards AITowards AIToneBench

Claude Opus 5 (max) vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 82.4.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
GLM-5
#43 Elo 1783 · 82.4/100
Cost / script
$0.120 vs $0.023
Human baseline
92.7 Claude Opus 5 (max) falls below it · GLM-5 falls below it

The verdict

GLM-5 is the cheapest serious option in this comparison at just over two cents a script, about a fifth of what Opus 5 costs, with open weights on top. At #43 and 82.42 overall against Opus at #1 and 90.80, you're trading eight points for that. Where it hurts most is anti-slop: 75.66 against 90.26, a fourteen point gap and the widest in this pairing. That's AI-flavoured filler and shaky number handling, and it's the kind of thing an editor has to catch line by line rather than fix in one pass. YouTube best practices is the other weak spot at 78.28. Its hooks are genuinely decent though, 86.47, so openings are usually salvageable. If your pipeline has a strong human edit at the end and volume matters more than polish, two cents a script is a real argument. If the draft needs to be close to shippable, the slop gap will cost you more in editing time than you saved.

Pick Opus 5 (max) if you want drafts that are close to clean, especially on slop and numbers.
Pick GLM-5 if cost dominates and you have real editorial capacity. Open weights at two cents a script, with decent hooks to build on.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
83.1
Writing Craft & Clarity13% weight
91.1
84.0
Substance, Accuracy & Value15% weight
90.6
84.0
Continuity & Emotion14% weight
90.4
79.7
YouTube Best Practices12% weight
89.6
78.3
Hook Strength10% weight
92.0
86.5
Length Adherence8% weight
89.7
78.6
Slop Score (EQ-Bench + ours)5% weight
93.3
85.3
Visual Cue Quality4% weight
89.6
83.8

Everything else that differs

Claude Opus 5 (max)GLM-5
Overall / 10090.882.4
Writing Elo26491783
Run-to-run spread (± overall std)1.2903.760
Cost per script (USD)0.1200.023
Avg latency (s)243.8117.3
Open weightsNoYes

Full scorecards: Claude Opus 5 (max) · GLM-5. How scoring works: methodology.

← All comparisons