Towards AITowards AIToneBench

Claude Opus 5 (max) vs Qwen3.7 Max (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 81.5.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
Cost / script
$0.120 vs $0.054
Human baseline
92.7 Claude Opus 5 (max) falls below it · Qwen3.7 Max (high) falls below it

The verdict

Turning Qwen3.7 Max up to high reasoning actually moves it down the board, #55 against the default's #47, 81.51 overall against 81.79. I expected the opposite. The high setting does buy slightly better hooks, 84.99 against the default's 84.25, and marginally better substance, but it gives back cue quality and a sliver of voice, and the board notices. It also costs the same, about five cents, so you're not even paying for the trade. Against Opus 5 at max effort it's nine points back, 90.80 to 81.51, with 2649.3 Elo against 1673.7. Where it stays weak is the same place the default is weak: continuity at 79.16 and anti-slop at 77.78, both roughly twelve points behind Opus. So the reasoning budget is not fixing the thing that's actually broken. If you want Qwen for this job, run the default and keep the three places.

Pick Opus 5 (max) if you want the gap closed rather than narrowed by a rounding error.
Pick Qwen3.7 Max (high) over the default config if you're using Qwen anyway. Same price, slightly better hooks.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
80.9
Writing Craft & Clarity13% weight
91.1
81.5
Substance, Accuracy & Value15% weight
90.6
81.2
Continuity & Emotion14% weight
90.4
79.2
YouTube Best Practices12% weight
89.6
80.7
Hook Strength10% weight
92.0
85.0
Length Adherence8% weight
89.7
83.3
Slop Score (EQ-Bench + ours)5% weight
93.3
85.9
Visual Cue Quality4% weight
89.6
78.4

Everything else that differs

Claude Opus 5 (max)Qwen3.7 Max (high)
Overall / 10090.881.5
Writing Elo26491674
Run-to-run spread (± overall std)1.2902.580
Cost per script (USD)0.1200.054
Avg latency (s)243.8157.2
Open weightsNoNo

Full scorecards: Claude Opus 5 (max) · Qwen3.7 Max (high). How scoring works: methodology.

← All comparisons