Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 82.4.
I wanted this one to be closer than it is, because GLM-5 is exactly the kind of open model I root for. But on our board Fable 5 wins comfortably: 90.0 to 82.4 overall, an Elo gap of nearly 800 points with confidence intervals nowhere near touching. When I read where GLM-5 actually loses it, two things stand out to me. Slop and number handling is the big one, 77.3 against Fable's 90.3, and with a standard deviation near 7 there I genuinely can't predict which GLM I'll get: some drafts come out clean, others read like a stats sheet with filler. Flow and emotion is the other gap, 79.7 vs 89, which on a long script is the difference between a story and a list of adjacent facts. The fair part: GLM-5 is open weights and costs about a cent and a half per script, estimated. If every draft gets a human rewrite anyway, that's a real argument. For publish-ready work, I'd pay Fable's roughly 75x premium.
Pick Fable 5 if you need consistent, low-slop scripts that hold their emotional thread across a full video.
Pick GLM-5 if you want open weights at roughly a cent and a half per draft and a human is rewriting everything anyway.
Blue bars: Claude Fable 5 (max). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | GLM-5 | |
|---|---|---|
| Overall / 100 | 90.3 | 82.4 |
| Writing Elo | 2575 | 1783 |
| Run-to-run spread (± overall std) | 1.610 | 3.760 |
| Cost per script (USD) | 0.908 | 0.023 |
| Avg latency (s) | 452.8 | 117.3 |
| Open weights | No | Yes |
Full scorecards: Claude Fable 5 (max) · GLM-5. How scoring works: methodology.
← All comparisons