Towards AITowards AIToneBench

Claude Fable 5 (max) vs Grok 4.6

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 86.6.

Claude Fable 5 (max)
#1 Elo 2321 · 89.3/100
Grok 4.6
#26 Elo 2035 · 86.6/100
Cost / script
$0.256 vs $0.204
Human baseline
90.2 Claude Fable 5 (max) falls below it · Grok 4.6 falls below it

The verdict

Fable 5 wins this one, and the intervals leave no doubt: Fable's Elo floor of 2272.7 sits well above Grok's ceiling of 2080.8, so the gap is statistically real. Fable is #1 on the board at 2320.9 Elo, Grok 4.6 way down at 2035.3, a 285.6-point spread, and overall it's 89.3 against 86.61. The widest gap is the one that matters most for retention: hook strength, 89.39 vs 80.81. Length adherence follows at 93.2 to 88.03. Grok takes exactly one metric, visual cues, 86.63 against Fable's 84.81. The records agree, 1185 wins and a single loss for Fable against Grok's 941 and 99. What Grok doesn't offer is a budget escape: at about twenty cents per script to Fable's roughly twenty-six, the saving is around five cents a draft. Grok does finish in less than half the time, which matters if you're iterating live. But if the script ships close to as-is, five cents doesn't cover a hook gap that wide.

Pick Claude Fable 5 (max effort + 4.8 fallback) if you want the top of the board, the strongest hooks and length adherence, and scripts that ship with minimal editing at roughly twenty-six cents each.
Pick Grok 4.6 if you want drafts in less than half the wait with slightly better visual cues at about twenty cents a script, and can live with a clear gap on hooks and overall quality.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: Grok 4.6. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.4
87.9
Writing Craft & Clarity13% weight
88.8
86.5
Substance, Accuracy & Value15% weight
88.5
87.5
Continuity & Emotion14% weight
88.3
85.8
YouTube Best Practices12% weight
89.0
86.5
Hook Strength10% weight
89.4
80.8
Length Adherence8% weight
93.2
88.0
Slop Score (EQ-Bench + ours)5% weight
93.3
91.2
Visual Cue Quality4% weight
84.8
86.6

Everything else that differs

Claude Fable 5 (max)Grok 4.6
Overall / 10089.386.6
Writing Elo23212035
Run-to-run spread (± overall std)1.5502.530
Cost per script (USD)0.2560.204
Avg latency (s)546.9243.1
Open weightsNoNo

Full scorecards: Claude Fable 5 (max) · Grok 4.6. How scoring works: methodology.

← All comparisons