Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 81.1.
This is about as one-sided as the board gets. Fable 5 sits at #1 with 2320.9 Elo; Qwen3.8 Max is way down the table at 1713.9, a 607-point gap, and the confidence intervals are nowhere close to touching, so the ranking is real. Overall it's 89.3 against 81.06, an 8.24-point gap, and Fable wins every single metric. The widest gaps are the ones that matter for a script: continuity and emotion, 88.26 vs 76.22, and YouTube best practices, 89.0 vs 77.76. Qwen is also erratic, with an overall standard deviation of 7.89 against Fable's 1.55, so some drafts land and others fall apart. The records repeat the story: 1185 wins and a single loss for Fable, 588 wins and 258 losses for Qwen. The usual consolation prize is cost, and here it's thin. About sixteen cents per script against about twenty-six. That discount doesn't come close to buying back the gap.
Pick Claude Fable 5 (max effort + 4.8 fallback) if you want the #1 board result, a 12.04-point edge in continuity and emotion, and drafts consistent enough to ship at about twenty-six cents each.
Pick Qwen3.8 Max if you only need rough first-pass drafts at about sixteen cents a script and can live with an 8.24-point overall gap and much swingier quality.
Blue bars: Claude Fable 5 (max). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 89.3 | 81.1 |
| Writing Elo | 2321 | 1714 |
| Run-to-run spread (± overall std) | 1.550 | 7.890 |
| Cost per script (USD) | 0.256 | 0.164 |
| Avg latency (s) | 546.9 | 429.0 |
| Open weights | No | No |
Full scorecards: Claude Fable 5 (max) · Qwen3.8 Max. How scoring works: methodology.
← All comparisons