Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 81.7.
Qwen3.7 Max sits at #50 with 81.73 overall against Opus 5's 90.80, so nine and a half points. Its strongest metric by some distance is length adherence at 85.79, which is respectable and better than several models ranked well above it. Hooks are decent too at 84.25. The problem is everything in the middle: continuity at 79.03, anti-slop at 77.84, tone and voice at 81.15. That combination reads as scripts that start well, hit the word count, and feel generic in between. Cost is about five and a half cents against Opus at twelve, so you're saving roughly half for a nine point drop. That's a worse trade than the genuinely cheap open models make, and it's closed weights, so you don't get the hosting argument either. Hard to see the case unless you're already committed to Qwen tooling.
Pick Opus 5 (max) if quality per dollar is the question. Twice the price for nine points is the better side of this trade.
Pick Qwen3.7 Max if you're already on Alibaba infrastructure and want half the cost with reliable length control.
Blue bars: Claude Opus 5 (max). Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | Qwen3.7 Max (default) | |
|---|---|---|
| Overall / 100 | 90.8 | 81.7 |
| Writing Elo | 2649 | 1699 |
| Run-to-run spread (± overall std) | 1.290 | 3.060 |
| Cost per script (USD) | 0.120 | 0.054 |
| Avg latency (s) | 243.8 | 156.8 |
| Open weights | No | No |
Full scorecards: Claude Opus 5 (max) · Qwen3.7 Max (default). How scoring works: methodology.
← All comparisons