Towards AITowards AIToneBench

Claude Fable 5 (max) vs DeepSeek V4 Pro (xhigh)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 80.5.

Claude Fable 5 (max)
#2 Elo 2575 · 90.3/100
DeepSeek V4 Pro (xhigh)
#60 Elo 1601 · 80.5/100
Cost / script
$0.908 vs $0.014
Human baseline
92.7 Claude Fable 5 (max) falls below it · DeepSeek V4 Pro (xhigh) falls below it

The verdict

This is the biggest gap across these seven match-ups, and it is not close. Claude Fable 5 at max effort sits second on my board, and the Elo distance here falls far outside both confidence intervals, so I do not need to hedge. Where I feel it when I read the scripts is the numbers. Fable pushes exact figures into on-screen cues the way I script them myself, scoring 90.12 on our numbers check, while DeepSeek V4 Pro recites stats into the spoken line and lands at 65.96, its weakest metric. Visual cues follow the same pattern. But cost is the honest complication. Fable runs about $1.10 per script and DeepSeek costs a cent and a half, roughly 80 times cheaper, with open weights you can run yourself. And credit where due: at that price, an 82 overall is genuinely solid work. If a human edits every draft anyway, DeepSeek is a defensible pipeline choice. If the script ships close to final, I pay for Fable.

Pick Claude Fable 5 (max effort) if the script ships close to final and the second-best writing on my board is worth about $1.10 a run.
Pick DeepSeek V4 Pro (xhigh) if a human edits every draft anyway and you want open weights at a cent and a half per script.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
90.4
82.2
Writing Craft & Clarity13% weight
89.6
82.2
Substance, Accuracy & Value15% weight
90.2
81.0
Continuity & Emotion14% weight
89.1
78.5
YouTube Best Practices12% weight
89.0
76.9
Hook Strength10% weight
90.1
84.6
Length Adherence8% weight
93.5
80.5
Slop Score (EQ-Bench + ours)5% weight
93.7
79.4
Visual Cue Quality4% weight
90.0
73.6

Everything else that differs

Claude Fable 5 (max)DeepSeek V4 Pro (xhigh)
Overall / 10090.380.5
Writing Elo25751601
Run-to-run spread (± overall std)1.6103.670
Cost per script (USD)0.9080.014
Avg latency (s)452.8163.2
Open weightsNoYes

Full scorecards: Claude Fable 5 (max) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.

← All comparisons