Towards AITowards AIToneBench

Claude Fable 5 (max) vs DeepSeek V4 Pro 0813 (max)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 84.0.

Claude Fable 5 (max)
#1 Elo 2321 · 89.3/100
DeepSeek V4 Pro 0813 (max)
#37 Elo 1881 · 84.0/100
Cost / script
$0.256 vs $0.022
Human baseline
90.2 Claude Fable 5 (max) falls below it · DeepSeek V4 Pro 0813 (max) falls below it

The verdict

Fable 5 wins this one, and it isn't close. Fable sits at #1 with 2320.9 Elo; DeepSeek V4 Pro is far down the board at 1880.7, a 440.2-point gap, and the confidence intervals are nowhere near touching, so the ranking is real. Overall it's 89.3 against 84.04, a 5.26-point gap, and Fable takes every individual metric. The widest gaps are the ones that decide whether a script is usable as written: YouTube best practices, 89.0 vs 81.13, continuity and emotion, 88.26 vs 80.77, and length adherence, 93.2 to 86.24. DeepSeek's best showing is hook strength, 86.83 against Fable's 89.39, still a loss. It's also less steady, with an overall std of 4.37 to Fable's 1.55, so a given draft can land well below its average. Then the budget line, and it's dramatic: DeepSeek is open weights at about two cents per script against roughly twenty-six cents for Fable, nearly 12x cheaper. That buys a lot of human editing, but here the quality gap is wide enough that the trade is hard to defend for anything shipping close to as-is.

Pick Claude Fable 5 (max effort + 4.8 fallback) if you want the #1 board result and its steadiness — every metric won, an overall std of 1.55 against 4.38 — at roughly twenty-six cents a script.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights at about two cents a script and can treat the 5.26-point overall gap as something your editing pass will absorb.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.4
84.7
Writing Craft & Clarity13% weight
88.8
84.9
Substance, Accuracy & Value15% weight
88.5
84.7
Continuity & Emotion14% weight
88.3
80.8
YouTube Best Practices12% weight
89.0
81.1
Hook Strength10% weight
89.4
86.8
Length Adherence8% weight
93.2
86.2
Slop Score (EQ-Bench + ours)5% weight
93.3
88.0
Visual Cue Quality4% weight
84.8
79.5

Everything else that differs

Claude Fable 5 (max)DeepSeek V4 Pro 0813 (max)
Overall / 10089.384.0
Writing Elo23211881
Run-to-run spread (± overall std)1.5504.370
Cost per script (USD)0.2560.022
Avg latency (s)546.9195.2
Open weightsNoYes

Full scorecards: Claude Fable 5 (max) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.

← All comparisons