Towards AITowards AIToneBench

Claude Fable 5 (max) vs GLM-5.3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 86.8.

Claude Fable 5 (max)
#1 Elo 2321 · 89.3/100
GLM-5.3
#24 Elo 2042 · 86.8/100
Cost / script
$0.256 vs $0.109
Human baseline
90.2 Claude Fable 5 (max) falls below it · GLM-5.3 falls below it

The verdict

Fable 5 wins this one, and the intervals say it's real: 2320.9 Elo at #1 against GLM-5.3's 2041.5 at #24, and the confidence bands don't come close to touching. Overall it's 89.3 to 86.83, a 2.47-point gap. GLM-5.3 actually takes one metric, hook strength, 89.54 to 89.39, but Fable wins everything else: length adherence is the widest gap at 93.2 against 87.5, YouTube best practices follows at 89.0 to 85.02, and continuity and emotion goes 88.26 to 85.11. The part the averages hide is consistency. Fable's overall spread is 1.55, GLM's is 6.6, so GLM swings from strong scripts to rough ones while Fable barely moves. The records tell the same story, 1185 wins and a single loss against 866 wins and 8 losses. Now the case for GLM: it's open weights, about eleven cents per script against roughly twenty-six for Fable, and it's noticeably faster. Staying within 2.47 points of the board leader while being self-hostable at under half the price is a real result. The inconsistency is what would worry me on anything that ships without an editing pass.

Pick Claude Fable 5 (max effort + 4.8 fallback) if you want the #1 board result with a tight 1.55-point spread, so every script lands close to its 89.3 average at roughly twenty-six cents each.
Pick GLM-5.3 if you want open weights you can self-host at about eleven cents per script and can live with output that swings between excellent and rough.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: GLM-5.3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.4
87.9
Writing Craft & Clarity13% weight
88.8
87.2
Substance, Accuracy & Value15% weight
88.5
85.9
Continuity & Emotion14% weight
88.3
85.1
YouTube Best Practices12% weight
89.0
85.0
Hook Strength10% weight
89.4
89.5
Length Adherence8% weight
93.2
87.5
Slop Score (EQ-Bench + ours)5% weight
93.3
91.6
Visual Cue Quality4% weight
84.8
81.4

Everything else that differs

Claude Fable 5 (max)GLM-5.3
Overall / 10089.386.8
Writing Elo23212042
Run-to-run spread (± overall std)1.5506.600
Cost per script (USD)0.2560.109
Avg latency (s)546.9301.9
Open weightsNoYes

Full scorecards: Claude Fable 5 (max) · GLM-5.3. How scoring works: methodology.

← All comparisons