Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 80.5.
This is the biggest gap across these seven match-ups, and it is not close. Claude Fable 5 at max effort sits second on my board, and the Elo distance here falls far outside both confidence intervals, so I do not need to hedge. Where I feel it when I read the scripts is the numbers. Fable pushes exact figures into on-screen cues the way I script them myself, scoring 90.12 on our numbers check, while DeepSeek V4 Pro recites stats into the spoken line and lands at 65.96, its weakest metric. Visual cues follow the same pattern. But cost is the honest complication. Fable runs about $1.10 per script and DeepSeek costs a cent and a half, roughly 80 times cheaper, with open weights you can run yourself. And credit where due: at that price, an 82 overall is genuinely solid work. If a human edits every draft anyway, DeepSeek is a defensible pipeline choice. If the script ships close to final, I pay for Fable.
Pick Claude Fable 5 (max effort) if the script ships close to final and the second-best writing on my board is worth about $1.10 a run.
Pick DeepSeek V4 Pro (xhigh) if a human edits every draft anyway and you want open weights at a cent and a half per script.
Blue bars: Claude Fable 5 (max). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | DeepSeek V4 Pro (xhigh) | |
|---|---|---|
| Overall / 100 | 90.3 | 80.5 |
| Writing Elo | 2575 | 1601 |
| Run-to-run spread (± overall std) | 1.610 | 3.670 |
| Cost per script (USD) | 0.908 | 0.014 |
| Avg latency (s) | 452.8 | 163.2 |
| Open weights | No | Yes |
Full scorecards: Claude Fable 5 (max) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.
← All comparisons