Towards AITowards AIToneBench

Claude Fable 5 (max) vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 81.1.

Claude Fable 5 (max)
#1 Elo 2321 · 89.3/100
Qwen3.8 Max
#46 Elo 1714 · 81.1/100
Cost / script
$0.256 vs $0.164
Human baseline
90.2 Claude Fable 5 (max) falls below it · Qwen3.8 Max falls below it

The verdict

This is about as one-sided as the board gets. Fable 5 sits at #1 with 2320.9 Elo; Qwen3.8 Max is way down the table at 1713.9, a 607-point gap, and the confidence intervals are nowhere close to touching, so the ranking is real. Overall it's 89.3 against 81.06, an 8.24-point gap, and Fable wins every single metric. The widest gaps are the ones that matter for a script: continuity and emotion, 88.26 vs 76.22, and YouTube best practices, 89.0 vs 77.76. Qwen is also erratic, with an overall standard deviation of 7.89 against Fable's 1.55, so some drafts land and others fall apart. The records repeat the story: 1185 wins and a single loss for Fable, 588 wins and 258 losses for Qwen. The usual consolation prize is cost, and here it's thin. About sixteen cents per script against about twenty-six. That discount doesn't come close to buying back the gap.

Pick Claude Fable 5 (max effort + 4.8 fallback) if you want the #1 board result, a 12.04-point edge in continuity and emotion, and drafts consistent enough to ship at about twenty-six cents each.
Pick Qwen3.8 Max if you only need rough first-pass drafts at about sixteen cents a script and can live with an 8.24-point overall gap and much swingier quality.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.4
80.7
Writing Craft & Clarity13% weight
88.8
81.0
Substance, Accuracy & Value15% weight
88.5
81.3
Continuity & Emotion14% weight
88.3
76.2
YouTube Best Practices12% weight
89.0
77.8
Hook Strength10% weight
89.4
83.4
Length Adherence8% weight
93.2
86.3
Slop Score (EQ-Bench + ours)5% weight
93.3
90.4
Visual Cue Quality4% weight
84.8
80.6

Everything else that differs

Claude Fable 5 (max)Qwen3.8 Max
Overall / 10089.381.1
Writing Elo23211714
Run-to-run spread (± overall std)1.5507.890
Cost per script (USD)0.2560.164
Avg latency (s)546.9429.0
Open weightsNoNo

Full scorecards: Claude Fable 5 (max) · Qwen3.8 Max. How scoring works: methodology.

← All comparisons