Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 81.5.
Turning Qwen3.7 Max up to high reasoning actually moves it down the board, #55 against the default's #47, 81.51 overall against 81.79. I expected the opposite. The high setting does buy slightly better hooks, 84.99 against the default's 84.25, and marginally better substance, but it gives back cue quality and a sliver of voice, and the board notices. It also costs the same, about five cents, so you're not even paying for the trade. Against Opus 5 at max effort it's nine points back, 90.80 to 81.51, with 2649.3 Elo against 1673.7. Where it stays weak is the same place the default is weak: continuity at 79.16 and anti-slop at 77.78, both roughly twelve points behind Opus. So the reasoning budget is not fixing the thing that's actually broken. If you want Qwen for this job, run the default and keep the three places.
Pick Opus 5 (max) if you want the gap closed rather than narrowed by a rounding error.
Pick Qwen3.7 Max (high) over the default config if you're using Qwen anyway. Same price, slightly better hooks.
Blue bars: Claude Opus 5 (max). Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | Qwen3.7 Max (high) | |
|---|---|---|
| Overall / 100 | 90.8 | 81.5 |
| Writing Elo | 2649 | 1674 |
| Run-to-run spread (± overall std) | 1.290 | 2.580 |
| Cost per script (USD) | 0.120 | 0.054 |
| Avg latency (s) | 243.8 | 157.2 |
| Open weights | No | No |
Full scorecards: Claude Opus 5 (max) · Qwen3.7 Max (high). How scoring works: methodology.
← All comparisons