Towards AITowards AIToneBench

Claude Opus 5 (max) vs Qwen3.7 Max (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 81.7.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
Qwen3.7 Max (default)
#50 Elo 1699 · 81.7/100
Cost / script
$0.120 vs $0.054
Human baseline
92.7 Claude Opus 5 (max) falls below it · Qwen3.7 Max (default) falls below it

The verdict

Qwen3.7 Max sits at #50 with 81.73 overall against Opus 5's 90.80, so nine and a half points. Its strongest metric by some distance is length adherence at 85.79, which is respectable and better than several models ranked well above it. Hooks are decent too at 84.25. The problem is everything in the middle: continuity at 79.03, anti-slop at 77.84, tone and voice at 81.15. That combination reads as scripts that start well, hit the word count, and feel generic in between. Cost is about five and a half cents against Opus at twelve, so you're saving roughly half for a nine point drop. That's a worse trade than the genuinely cheap open models make, and it's closed weights, so you don't get the hosting argument either. Hard to see the case unless you're already committed to Qwen tooling.

Pick Opus 5 (max) if quality per dollar is the question. Twice the price for nine points is the better side of this trade.
Pick Qwen3.7 Max if you're already on Alibaba infrastructure and want half the cost with reliable length control.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
81.2
Writing Craft & Clarity13% weight
91.1
81.7
Substance, Accuracy & Value15% weight
90.6
80.9
Continuity & Emotion14% weight
90.4
79.0
YouTube Best Practices12% weight
89.6
80.8
Hook Strength10% weight
92.0
84.2
Length Adherence8% weight
89.7
85.8
Slop Score (EQ-Bench + ours)5% weight
93.3
86.0
Visual Cue Quality4% weight
89.6
80.2

Everything else that differs

Claude Opus 5 (max)Qwen3.7 Max (default)
Overall / 10090.881.7
Writing Elo26491699
Run-to-run spread (± overall std)1.2903.060
Cost per script (USD)0.1200.054
Avg latency (s)243.8156.8
Open weightsNoNo

Full scorecards: Claude Opus 5 (max) · Qwen3.7 Max (default). How scoring works: methodology.

← All comparisons