Towards AITowards AIToneBench

GPT-5.6 Sol (ultra) vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.4 to 83.4.

GPT-5.6 Sol (ultra)
#16 Elo 2316 · 88.4/100
Qwen3.8 Max
#38 Elo 1866 · 83.4/100
Cost / script
$0.137 vs $0.150
Human baseline
92.7 GPT-5.6 Sol (ultra) falls below it · Qwen3.8 Max falls below it

The verdict

Both of these are length specialists, which makes the comparison unusually clean. GPT-5.6 Sol at ultra is #16 with 94.34 length adherence; Qwen3.8 Max is #38 with 92.90. Under a point apart on the metric they both do best. Everything else favours GPT, and the overall is 88.40 against 83.44. The decisive gap is continuity: 85.72 against 78.08, nearly eight points. GPT also takes substance 90.49 to 84.37 and cue quality 91.17 to 85.22. Where Qwen stays competitive is anti-slop at 87.78 against 90.60, so it is not a slop machine, it just does not hold a long script together. Cost is close, 14 cents against 15, and both are slow, 421 seconds against 397. Given near-identical price and speed, and five points of quality, this one is not close on the merits.

Pick GPT-5.6 Sol (ultra). Same price, same speed, five points better, and the flow gap is the one you feel reading it.
Pick Qwen3.8 Max only if you are already on Alibaba infrastructure and want comparable length discipline there.

Metric by metric

Blue bars: GPT-5.6 Sol (ultra). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.0
81.9
Writing Craft & Clarity13% weight
88.1
83.4
Substance, Accuracy & Value15% weight
90.5
84.4
Continuity & Emotion14% weight
85.7
78.1
YouTube Best Practices12% weight
86.1
80.4
Hook Strength10% weight
86.4
83.5
Length Adherence8% weight
94.3
92.9
Slop Score (EQ-Bench + ours)5% weight
93.5
92.0
Visual Cue Quality4% weight
91.2
85.2

Everything else that differs

GPT-5.6 Sol (ultra)Qwen3.8 Max
Overall / 10088.483.4
Writing Elo23161866
Run-to-run spread (± overall std)1.9705.620
Cost per script (USD)0.1370.150
Avg latency (s)420.9397.4
Open weightsNoNo

Full scorecards: GPT-5.6 Sol (ultra) · Qwen3.8 Max. How scoring works: methodology.

← All comparisons