Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 81.5.
Fable 5 takes this one without much drama: 90.21 to 81.51 overall, with Elo intervals separated by hundreds of points, so I don't need to hedge. The gap is broad rather than one dramatic failure. Qwen3.7 Max trails on tone and voice, 80.91 vs 90.41, and for this benchmark that's the whole point: scripts are judged blind against my own reference scripts, and Qwen doesn't sound like a person telling you something. Its visual cues sit among its weakest metrics at 78.42, and they're wobbly on top of that, so plan to redo the on-screen work yourself. Where Qwen holds its own: hooks at 84.99 are respectable, and its length discipline is decent. It's also about 27x cheaper at roughly five cents per script, though that pricing is estimated. My honest take: at this quality distance the cost argument gets weaker, because you pay the difference back in editing time. And if budget rules everything, cheaper models on this board outscore Qwen anyway.
Pick Fable 5 if the script needs to sound like a human voice on camera, not a well-formatted document.
Pick Qwen3.7 Max if you need a budget drafter for structure and hooks and you'll rewrite the voice and cues yourself.
Blue bars: Claude Fable 5 (max). Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | Qwen3.7 Max (high) | |
|---|---|---|
| Overall / 100 | 90.3 | 81.5 |
| Writing Elo | 2575 | 1674 |
| Run-to-run spread (± overall std) | 1.610 | 2.580 |
| Cost per script (USD) | 0.908 | 0.054 |
| Avg latency (s) | 452.8 | 157.2 |
| Open weights | No | No |
Full scorecards: Claude Fable 5 (max) · Qwen3.7 Max (high). How scoring works: methodology.
← All comparisons