Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 85.8.
No coin flip here. Fable 5 sits about 530 Elo points above Grok 4.5 and the confidence intervals aren't anywhere near each other. Two metrics tell the story: length adherence, where Grok scores 78.67 against Fable's 93.49, and continuity and emotion, 83.5 vs 88.95. Translated: Grok writes decent sentences but drifts off the target length and loses the emotional thread that keeps someone watching. Credit where it's due, though. Grok's hooks are genuinely solid at 87.96, and it's absurdly fast, under 30 seconds per script. And then the price: roughly two cents per script versus $1.10. That's around 65x cheaper, which is not a rounding error. I wouldn't publish a Grok draft as-is on my channel, but as a brainstorming pass or a hook generator you rewrite afterward, two cents buys a lot of raw material. For a script that has to hold attention for ten straight minutes, Fable is the one doing the actual job.
Pick Fable 5 if the script needs to hold its length and emotional flow well enough to publish with light edits.
Pick Grok 4.5 if you want near-instant, two-cent drafts for ideation and hooks, and you plan to rewrite the rest anyway.
Blue bars: Claude Fable 5 (max). Orange bars: Grok 4.5. Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | Grok 4.5 | |
|---|---|---|
| Overall / 100 | 90.3 | 85.8 |
| Writing Elo | 2575 | 2042 |
| Run-to-run spread (± overall std) | 1.610 | 2.810 |
| Cost per script (USD) | 0.908 | 0.038 |
| Avg latency (s) | 452.8 | 41.7 |
| Open weights | No | No |
Full scorecards: Claude Fable 5 (max) · Grok 4.5. How scoring works: methodology.
← All comparisons