Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 83.4.
Qwen3.8 Max is a real jump over 3.7: #38 against #47, 83.44 overall against 81.79. Two points and fifteen places. Against Opus 5 at max effort the gap is still 7 points, 90.80 to 83.44, with 2649.3 Elo against 1865.6. The shape is lopsided in an interesting way. Qwen's length adherence is 92.90 against Opus 5's 89.73, so it actually beats the board leader on hitting a target word count, and its anti-slop score of 87.78 is respectable. Where it falls apart is continuity at 78.08 against 90.50. Twelve and a half points. It writes well-sized, clean-ish paragraphs that do not hold together as one piece across a long script. Voice is the other soft spot, 81.93 against 91.22. Cost is the surprise: 15 cents a script against Opus 5's 12, so the cheaper-looking model is the more expensive one here, and it is three times slower at 397 seconds. Hard to justify unless you are already on Alibaba tooling.
Pick Opus 5 (max) on quality per dollar. It is cheaper, faster and seven points better.
Pick Qwen3.8 Max if you need its length discipline specifically, or you are already committed to Alibaba infrastructure.
Blue bars: Claude Opus 5 (max). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 90.8 | 83.4 |
| Writing Elo | 2649 | 1866 |
| Run-to-run spread (± overall std) | 1.290 | 5.620 |
| Cost per script (USD) | 0.120 | 0.150 |
| Avg latency (s) | 243.8 | 397.4 |
| Open weights | No | No |
Full scorecards: Claude Opus 5 (max) · Qwen3.8 Max. How scoring works: methodology.
← All comparisons