Towards AITowards AIToneBench

Claude Opus 5 (max) vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 83.4.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
Qwen3.8 Max
#38 Elo 1866 · 83.4/100
Cost / script
$0.120 vs $0.150
Human baseline
92.7 Claude Opus 5 (max) falls below it · Qwen3.8 Max falls below it

The verdict

Qwen3.8 Max is a real jump over 3.7: #38 against #47, 83.44 overall against 81.79. Two points and fifteen places. Against Opus 5 at max effort the gap is still 7 points, 90.80 to 83.44, with 2649.3 Elo against 1865.6. The shape is lopsided in an interesting way. Qwen's length adherence is 92.90 against Opus 5's 89.73, so it actually beats the board leader on hitting a target word count, and its anti-slop score of 87.78 is respectable. Where it falls apart is continuity at 78.08 against 90.50. Twelve and a half points. It writes well-sized, clean-ish paragraphs that do not hold together as one piece across a long script. Voice is the other soft spot, 81.93 against 91.22. Cost is the surprise: 15 cents a script against Opus 5's 12, so the cheaper-looking model is the more expensive one here, and it is three times slower at 397 seconds. Hard to justify unless you are already on Alibaba tooling.

Pick Opus 5 (max) on quality per dollar. It is cheaper, faster and seven points better.
Pick Qwen3.8 Max if you need its length discipline specifically, or you are already committed to Alibaba infrastructure.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
81.9
Writing Craft & Clarity13% weight
91.1
83.4
Substance, Accuracy & Value15% weight
90.6
84.4
Continuity & Emotion14% weight
90.4
78.1
YouTube Best Practices12% weight
89.6
80.4
Hook Strength10% weight
92.0
83.5
Length Adherence8% weight
89.7
92.9
Slop Score (EQ-Bench + ours)5% weight
93.3
92.0
Visual Cue Quality4% weight
89.6
85.2

Everything else that differs

Claude Opus 5 (max)Qwen3.8 Max
Overall / 10090.883.4
Writing Elo26491866
Run-to-run spread (± overall std)1.2905.620
Cost per script (USD)0.1200.150
Avg latency (s)243.8397.4
Open weightsNoNo

Full scorecards: Claude Opus 5 (max) · Qwen3.8 Max. How scoring works: methodology.

← All comparisons