Towards AITowards AIToneBench

Claude Fable 5 (max) vs GPT-5.6 Sol (ultra)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 88.0.

Claude Fable 5 (max)
#1 Elo 2321 · 89.3/100
GPT-5.6 Sol (ultra)
#12 Elo 2177 · 88.0/100
Cost / script
$0.256 vs $0.143
Human baseline
90.2 Claude Fable 5 (max) falls below it · GPT-5.6 Sol (ultra) falls below it

The verdict

Fable 5 wins this one, and the intervals leave no doubt: its Elo floor of 2272.7 clears GPT-5.6 Ultra's ceiling of 2214.9. On the board it's #1 at 2320.9 against #12 at 2177.4, a 143.5-point Elo gap, and 89.30 to 87.97 overall. The gap lives in the writing itself. Tone and voice is 89.41 vs 86.80, continuity and emotion 88.26 vs 85.85, the two metrics that decide whether a script sounds like me or like a summary of me. GPT-5.6 wins where its effort seems to go: visual cues at 88.55 against Fable's 84.81, a 3.74-point edge, plus a slight lead on substance, 89.01 to 88.49. The records agree with the ranking, Fable at 1185 wins and a single loss against GPT-5.6's 1042 and ten. Then the bill: about fourteen cents per script against about twenty-six, so GPT-5.6 is still roughly half the price even at ultra. Same trade as always. If the voice is the product, pay for Fable. If the structure is what you need and the voice gets rewritten anyway, the cheaper draft is fine.

Pick Claude Fable 5 (max effort + 4.8 fallback) if the script ships under your name and the clear lead in tone, voice, and continuity is worth about twenty-six cents a draft.
Pick GPT-5.6 Sol (ultra) if you want the strongest visual cues in this pairing and a slight substance edge at about fourteen cents a script, roughly half Fable's price.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.4
86.8
Writing Craft & Clarity13% weight
88.8
87.7
Substance, Accuracy & Value15% weight
88.5
89.0
Continuity & Emotion14% weight
88.3
85.8
YouTube Best Practices12% weight
89.0
86.5
Hook Strength10% weight
89.4
87.7
Length Adherence8% weight
93.2
91.9
Slop Score (EQ-Bench + ours)5% weight
93.3
93.3
Visual Cue Quality4% weight
84.8
88.5

Everything else that differs

Claude Fable 5 (max)GPT-5.6 Sol (ultra)
Overall / 10089.388.0
Writing Elo23212177
Run-to-run spread (± overall std)1.5501.450
Cost per script (USD)0.2560.143
Avg latency (s)546.9195.5
Open weightsNoNo

Full scorecards: Claude Fable 5 (max) · GPT-5.6 Sol (ultra). How scoring works: methodology.

← All comparisons