Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 86.6.
Fable 5 wins this one, and the intervals leave no doubt: Fable's Elo floor of 2272.7 sits well above Grok's ceiling of 2080.8, so the gap is statistically real. Fable is #1 on the board at 2320.9 Elo, Grok 4.6 way down at 2035.3, a 285.6-point spread, and overall it's 89.3 against 86.61. The widest gap is the one that matters most for retention: hook strength, 89.39 vs 80.81. Length adherence follows at 93.2 to 88.03. Grok takes exactly one metric, visual cues, 86.63 against Fable's 84.81. The records agree, 1185 wins and a single loss for Fable against Grok's 941 and 99. What Grok doesn't offer is a budget escape: at about twenty cents per script to Fable's roughly twenty-six, the saving is around five cents a draft. Grok does finish in less than half the time, which matters if you're iterating live. But if the script ships close to as-is, five cents doesn't cover a hook gap that wide.
Pick Claude Fable 5 (max effort + 4.8 fallback) if you want the top of the board, the strongest hooks and length adherence, and scripts that ship with minimal editing at roughly twenty-six cents each.
Pick Grok 4.6 if you want drafts in less than half the wait with slightly better visual cues at about twenty cents a script, and can live with a clear gap on hooks and overall quality.
Blue bars: Claude Fable 5 (max). Orange bars: Grok 4.6. Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | Grok 4.6 | |
|---|---|---|
| Overall / 100 | 89.3 | 86.6 |
| Writing Elo | 2321 | 2035 |
| Run-to-run spread (± overall std) | 1.550 | 2.530 |
| Cost per script (USD) | 0.256 | 0.204 |
| Avg latency (s) | 546.9 | 243.1 |
| Open weights | No | No |
Full scorecards: Claude Fable 5 (max) · Grok 4.6. How scoring works: methodology.
← All comparisons