Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 82.4.
The headline gap is big, 90.21 vs 81.8 overall, but the number that actually decides this matchup is the variance. Fable 5's overall score moves about a point run to run. MiniMax M3's swings by about six, and its per-task scores range from 75.5 on my AI-engineering roadmap task up to about 87 on the deployment one. Its length adherence deviation alone is nearly 13 points. In plain terms: sometimes MiniMax hands you a decent script, sometimes a mess, and you won't know which until you read it. YouTube structure is its other weak spot, 77.7 vs Fable's 88.95, so retention beats and CTAs need rebuilding by hand. On the plus side, it's open weights and around two cents per script, estimated, and its hooks hold up fine at 86.1. But a writing tool you can't trust to be the same tool twice is a hard sell for anything with a deadline. Fable costs sixty times more and earns it here.
Pick Fable 5 if you need the same quality every run; consistency is the widest gap in this matchup.
Pick MiniMax M3 if two-cent open-weights drafts are worth a quality lottery you'll filter by hand.
Blue bars: Claude Fable 5 (max). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 90.3 | 82.4 |
| Writing Elo | 2575 | 1790 |
| Run-to-run spread (± overall std) | 1.610 | 5.610 |
| Cost per script (USD) | 0.908 | 0.019 |
| Avg latency (s) | 452.8 | 185.4 |
| Open weights | No | Yes |
Full scorecards: Claude Fable 5 (max) · MiniMax M3. How scoring works: methodology.
← All comparisons