Towards AITowards AIToneBench

Claude Fable 5 (max) vs Qwen3.7 Max (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 81.5.

Claude Fable 5 (max)
#2 Elo 2575 · 90.3/100
Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
Cost / script
$0.908 vs $0.054
Human baseline
92.7 Claude Fable 5 (max) falls below it · Qwen3.7 Max (high) falls below it

The verdict

Fable 5 takes this one without much drama: 90.21 to 81.51 overall, with Elo intervals separated by hundreds of points, so I don't need to hedge. The gap is broad rather than one dramatic failure. Qwen3.7 Max trails on tone and voice, 80.91 vs 90.41, and for this benchmark that's the whole point: scripts are judged blind against my own reference scripts, and Qwen doesn't sound like a person telling you something. Its visual cues sit among its weakest metrics at 78.42, and they're wobbly on top of that, so plan to redo the on-screen work yourself. Where Qwen holds its own: hooks at 84.99 are respectable, and its length discipline is decent. It's also about 27x cheaper at roughly five cents per script, though that pricing is estimated. My honest take: at this quality distance the cost argument gets weaker, because you pay the difference back in editing time. And if budget rules everything, cheaper models on this board outscore Qwen anyway.

Pick Fable 5 if the script needs to sound like a human voice on camera, not a well-formatted document.
Pick Qwen3.7 Max if you need a budget drafter for structure and hooks and you'll rewrite the voice and cues yourself.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
90.4
80.9
Writing Craft & Clarity13% weight
89.6
81.5
Substance, Accuracy & Value15% weight
90.2
81.2
Continuity & Emotion14% weight
89.1
79.2
YouTube Best Practices12% weight
89.0
80.7
Hook Strength10% weight
90.1
85.0
Length Adherence8% weight
93.5
83.3
Slop Score (EQ-Bench + ours)5% weight
93.7
85.9
Visual Cue Quality4% weight
90.0
78.4

Everything else that differs

Claude Fable 5 (max)Qwen3.7 Max (high)
Overall / 10090.381.5
Writing Elo25751674
Run-to-run spread (± overall std)1.6102.580
Cost per script (USD)0.9080.054
Avg latency (s)452.8157.2
Open weightsNoNo

Full scorecards: Claude Fable 5 (max) · Qwen3.7 Max (high). How scoring works: methodology.

← All comparisons