Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 82.4.
GLM-5 is the cheapest serious option in this comparison at just over two cents a script, about a fifth of what Opus 5 costs, with open weights on top. At #43 and 82.42 overall against Opus at #1 and 90.80, you're trading eight points for that. Where it hurts most is anti-slop: 75.66 against 90.26, a fourteen point gap and the widest in this pairing. That's AI-flavoured filler and shaky number handling, and it's the kind of thing an editor has to catch line by line rather than fix in one pass. YouTube best practices is the other weak spot at 78.28. Its hooks are genuinely decent though, 86.47, so openings are usually salvageable. If your pipeline has a strong human edit at the end and volume matters more than polish, two cents a script is a real argument. If the draft needs to be close to shippable, the slop gap will cost you more in editing time than you saved.
Pick Opus 5 (max) if you want drafts that are close to clean, especially on slop and numbers.
Pick GLM-5 if cost dominates and you have real editorial capacity. Open weights at two cents a script, with decent hooks to build on.
Blue bars: Claude Opus 5 (max). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | GLM-5 | |
|---|---|---|
| Overall / 100 | 90.8 | 82.4 |
| Writing Elo | 2649 | 1783 |
| Run-to-run spread (± overall std) | 1.290 | 3.760 |
| Cost per script (USD) | 0.120 | 0.023 |
| Avg latency (s) | 243.8 | 117.3 |
| Open weights | No | Yes |
Full scorecards: Claude Opus 5 (max) · GLM-5. How scoring works: methodology.
← All comparisons