Towards AITowards AIToneBench

Claude Fable 5 (max) vs GPT-5.6 Sol (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 88.4.

Claude Fable 5 (max)
#2 Elo 2575 · 90.3/100
GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
Cost / script
$0.908 vs $0.163
Human baseline
92.7 Claude Fable 5 (max) falls below it · GPT-5.6 Sol (high) falls below it

The verdict

This one is not close on the writing, and the Elo intervals don't overlap, so I'll say it plainly: Fable 5 wins as a writer. The gap shows up exactly where I care most. Tone and voice, 90.41 vs 87.31, and continuity and emotion, 88.95 vs 86.1. In practice, that means Fable sounds more like me telling a story and less like a well-organized report. GPT-5.6 fights back where it matters for production: it edges Fable on visual cues, 91.02 vs 90.00, ties it on substance, and writes a script in under two minutes at roughly an eighth of the price. Fifteen cents versus $1.10 per script adds up fast if you draft daily. So the honest read: Fable is the better writer, GPT-5.6 is the better deal, and neither reaches the 92.74 human baseline yet. If the final draft goes out under your name, the voice gap is worth paying for. If it's a first pass you'll rewrite anyway, it probably isn't.

Pick Fable 5 if the script ships close to as-written and voice fidelity is the thing you're paying for.
Pick GPT-5.6 Sol if you draft in volume: it's about 8x cheaper, much faster, and its visual cues are the best of the pair.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: GPT-5.6 Sol (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
90.4
87.3
Writing Craft & Clarity13% weight
89.6
88.1
Substance, Accuracy & Value15% weight
90.2
90.1
Continuity & Emotion14% weight
89.1
86.1
YouTube Best Practices12% weight
89.0
86.5
Hook Strength10% weight
90.1
87.5
Length Adherence8% weight
93.5
91.7
Slop Score (EQ-Bench + ours)5% weight
93.7
93.4
Visual Cue Quality4% weight
90.0
91.0

Everything else that differs

Claude Fable 5 (max)GPT-5.6 Sol (high)
Overall / 10090.388.4
Writing Elo25752315
Run-to-run spread (± overall std)1.6101.750
Cost per script (USD)0.9080.163
Avg latency (s)452.8106.4
Open weightsNoNo

Full scorecards: Claude Fable 5 (max) · GPT-5.6 Sol (high). How scoring works: methodology.

← All comparisons