Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 88.4.
This one is not close on the writing, and the Elo intervals don't overlap, so I'll say it plainly: Fable 5 wins as a writer. The gap shows up exactly where I care most. Tone and voice, 90.41 vs 87.31, and continuity and emotion, 88.95 vs 86.1. In practice, that means Fable sounds more like me telling a story and less like a well-organized report. GPT-5.6 fights back where it matters for production: it edges Fable on visual cues, 91.02 vs 90.00, ties it on substance, and writes a script in under two minutes at roughly an eighth of the price. Fifteen cents versus $1.10 per script adds up fast if you draft daily. So the honest read: Fable is the better writer, GPT-5.6 is the better deal, and neither reaches the 92.74 human baseline yet. If the final draft goes out under your name, the voice gap is worth paying for. If it's a first pass you'll rewrite anyway, it probably isn't.
Pick Fable 5 if the script ships close to as-written and voice fidelity is the thing you're paying for.
Pick GPT-5.6 Sol if you draft in volume: it's about 8x cheaper, much faster, and its visual cues are the best of the pair.
Blue bars: Claude Fable 5 (max). Orange bars: GPT-5.6 Sol (high). Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | GPT-5.6 Sol (high) | |
|---|---|---|
| Overall / 100 | 90.3 | 88.4 |
| Writing Elo | 2575 | 2315 |
| Run-to-run spread (± overall std) | 1.610 | 1.750 |
| Cost per script (USD) | 0.908 | 0.163 |
| Avg latency (s) | 452.8 | 106.4 |
| Open weights | No | No |
Full scorecards: Claude Fable 5 (max) · GPT-5.6 Sol (high). How scoring works: methodology.
← All comparisons