Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 84.0.
Fable 5 wins this one, and it isn't close. Fable sits at #1 with 2320.9 Elo; DeepSeek V4 Pro is far down the board at 1880.7, a 440.2-point gap, and the confidence intervals are nowhere near touching, so the ranking is real. Overall it's 89.3 against 84.04, a 5.26-point gap, and Fable takes every individual metric. The widest gaps are the ones that decide whether a script is usable as written: YouTube best practices, 89.0 vs 81.13, continuity and emotion, 88.26 vs 80.77, and length adherence, 93.2 to 86.24. DeepSeek's best showing is hook strength, 86.83 against Fable's 89.39, still a loss. It's also less steady, with an overall std of 4.37 to Fable's 1.55, so a given draft can land well below its average. Then the budget line, and it's dramatic: DeepSeek is open weights at about two cents per script against roughly twenty-six cents for Fable, nearly 12x cheaper. That buys a lot of human editing, but here the quality gap is wide enough that the trade is hard to defend for anything shipping close to as-is.
Pick Claude Fable 5 (max effort + 4.8 fallback) if you want the #1 board result and its steadiness — every metric won, an overall std of 1.55 against 4.38 — at roughly twenty-six cents a script.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights at about two cents a script and can treat the 5.26-point overall gap as something your editing pass will absorb.
Blue bars: Claude Fable 5 (max). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | DeepSeek V4 Pro 0813 (max) | |
|---|---|---|
| Overall / 100 | 89.3 | 84.0 |
| Writing Elo | 2321 | 1881 |
| Run-to-run spread (± overall std) | 1.550 | 4.370 |
| Cost per script (USD) | 0.256 | 0.022 |
| Avg latency (s) | 546.9 | 195.2 |
| Open weights | No | Yes |
Full scorecards: Claude Fable 5 (max) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.
← All comparisons