Towards AITowards AIToneBench

Claude Fable 5 (max) vs Grok 4.5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 85.8.

Claude Fable 5 (max)
#2 Elo 2575 · 90.3/100
Grok 4.5
#30 Elo 2042 · 85.8/100
Cost / script
$0.908 vs $0.038
Human baseline
92.7 Claude Fable 5 (max) falls below it · Grok 4.5 falls below it

The verdict

No coin flip here. Fable 5 sits about 530 Elo points above Grok 4.5 and the confidence intervals aren't anywhere near each other. Two metrics tell the story: length adherence, where Grok scores 78.67 against Fable's 93.49, and continuity and emotion, 83.5 vs 88.95. Translated: Grok writes decent sentences but drifts off the target length and loses the emotional thread that keeps someone watching. Credit where it's due, though. Grok's hooks are genuinely solid at 87.96, and it's absurdly fast, under 30 seconds per script. And then the price: roughly two cents per script versus $1.10. That's around 65x cheaper, which is not a rounding error. I wouldn't publish a Grok draft as-is on my channel, but as a brainstorming pass or a hook generator you rewrite afterward, two cents buys a lot of raw material. For a script that has to hold attention for ten straight minutes, Fable is the one doing the actual job.

Pick Fable 5 if the script needs to hold its length and emotional flow well enough to publish with light edits.
Pick Grok 4.5 if you want near-instant, two-cent drafts for ideation and hooks, and you plan to rewrite the rest anyway.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
90.4
87.1
Writing Craft & Clarity13% weight
89.6
86.6
Substance, Accuracy & Value15% weight
90.2
87.4
Continuity & Emotion14% weight
89.1
83.5
YouTube Best Practices12% weight
89.0
84.4
Hook Strength10% weight
90.1
88.0
Length Adherence8% weight
93.5
78.7
Slop Score (EQ-Bench + ours)5% weight
93.7
90.2
Visual Cue Quality4% weight
90.0
86.5

Everything else that differs

Claude Fable 5 (max)Grok 4.5
Overall / 10090.385.8
Writing Elo25752042
Run-to-run spread (± overall std)1.6102.810
Cost per script (USD)0.9080.038
Avg latency (s)452.841.7
Open weightsNoNo

Full scorecards: Claude Fable 5 (max) · Grok 4.5. How scoring works: methodology.

← All comparisons