Towards AITowards AIToneBench

Claude Fable 5 (max) vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 82.4.

Claude Fable 5 (max)
#2 Elo 2575 · 90.3/100
GLM-5
#43 Elo 1783 · 82.4/100
Cost / script
$0.908 vs $0.023
Human baseline
92.7 Claude Fable 5 (max) falls below it · GLM-5 falls below it

The verdict

I wanted this one to be closer than it is, because GLM-5 is exactly the kind of open model I root for. But on our board Fable 5 wins comfortably: 90.0 to 82.4 overall, an Elo gap of nearly 800 points with confidence intervals nowhere near touching. When I read where GLM-5 actually loses it, two things stand out to me. Slop and number handling is the big one, 77.3 against Fable's 90.3, and with a standard deviation near 7 there I genuinely can't predict which GLM I'll get: some drafts come out clean, others read like a stats sheet with filler. Flow and emotion is the other gap, 79.7 vs 89, which on a long script is the difference between a story and a list of adjacent facts. The fair part: GLM-5 is open weights and costs about a cent and a half per script, estimated. If every draft gets a human rewrite anyway, that's a real argument. For publish-ready work, I'd pay Fable's roughly 75x premium.

Pick Fable 5 if you need consistent, low-slop scripts that hold their emotional thread across a full video.
Pick GLM-5 if you want open weights at roughly a cent and a half per draft and a human is rewriting everything anyway.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
90.4
83.1
Writing Craft & Clarity13% weight
89.6
84.0
Substance, Accuracy & Value15% weight
90.2
84.0
Continuity & Emotion14% weight
89.1
79.7
YouTube Best Practices12% weight
89.0
78.3
Hook Strength10% weight
90.1
86.5
Length Adherence8% weight
93.5
78.6
Slop Score (EQ-Bench + ours)5% weight
93.7
85.3
Visual Cue Quality4% weight
90.0
83.8

Everything else that differs

Claude Fable 5 (max)GLM-5
Overall / 10090.382.4
Writing Elo25751783
Run-to-run spread (± overall std)1.6103.760
Cost per script (USD)0.9080.023
Avg latency (s)452.8117.3
Open weightsNoYes

Full scorecards: Claude Fable 5 (max) · GLM-5. How scoring works: methodology.

← All comparisons