Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 88.0.
Fable 5 wins this one, and the intervals leave no doubt: its Elo floor of 2272.7 clears GPT-5.6 Ultra's ceiling of 2214.9. On the board it's #1 at 2320.9 against #12 at 2177.4, a 143.5-point Elo gap, and 89.30 to 87.97 overall. The gap lives in the writing itself. Tone and voice is 89.41 vs 86.80, continuity and emotion 88.26 vs 85.85, the two metrics that decide whether a script sounds like me or like a summary of me. GPT-5.6 wins where its effort seems to go: visual cues at 88.55 against Fable's 84.81, a 3.74-point edge, plus a slight lead on substance, 89.01 to 88.49. The records agree with the ranking, Fable at 1185 wins and a single loss against GPT-5.6's 1042 and ten. Then the bill: about fourteen cents per script against about twenty-six, so GPT-5.6 is still roughly half the price even at ultra. Same trade as always. If the voice is the product, pay for Fable. If the structure is what you need and the voice gets rewritten anyway, the cheaper draft is fine.
Pick Claude Fable 5 (max effort + 4.8 fallback) if the script ships under your name and the clear lead in tone, voice, and continuity is worth about twenty-six cents a draft.
Pick GPT-5.6 Sol (ultra) if you want the strongest visual cues in this pairing and a slight substance edge at about fourteen cents a script, roughly half Fable's price.
Blue bars: Claude Fable 5 (max). Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | GPT-5.6 Sol (ultra) | |
|---|---|---|
| Overall / 100 | 89.3 | 88.0 |
| Writing Elo | 2321 | 2177 |
| Run-to-run spread (± overall std) | 1.550 | 1.450 |
| Cost per script (USD) | 0.256 | 0.143 |
| Avg latency (s) | 546.9 | 195.5 |
| Open weights | No | No |
Full scorecards: Claude Fable 5 (max) · GPT-5.6 Sol (ultra). How scoring works: methodology.
← All comparisons