Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs Qwen3.7 Max (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 81.7.

GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
Qwen3.7 Max (default)
#50 Elo 1699 · 81.7/100
Cost / script
$0.163 vs $0.054
Human baseline
92.7 GPT-5.6 Sol (high) falls below it · Qwen3.7 Max (default) falls below it

The verdict

GPT-5.6 Sol at high reasoning is #17, Qwen3.7 Max is #50, and the overall is 88.38 against 81.73. Just under seven points. The gap is widest on the mechanical metrics, which is GPT's whole personality on this board: anti-slop and numbers 90.60 against 77.84, cue quality 91.02 against 80.24. Both over twelve points. If your scripts are full of figures and visual direction, that's the entire argument. Continuity is the third big one, 86.07 to 79.03. Qwen's best number is length adherence at 85.79, and it's genuinely close to GPT's 91.69 by the standards of the rest of this comparison. On price and speed GPT is 16 cents and 106 seconds against Qwen's five cents and 157 seconds. So GPT is three times the price and noticeably faster, which is an unusual combination and worth knowing.

Pick GPT-5.6 Sol (high) if numbers and visual cues have to be right, and you want drafts back fast.
Pick Qwen3.7 Max if cost is the driver and you can absorb the slop and cue cleanup.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.3
81.2
Writing Craft & Clarity13% weight
88.1
81.7
Substance, Accuracy & Value15% weight
90.1
80.9
Continuity & Emotion14% weight
86.1
79.0
YouTube Best Practices12% weight
86.5
80.8
Hook Strength10% weight
87.5
84.2
Length Adherence8% weight
91.7
85.8
Slop Score (EQ-Bench + ours)5% weight
93.4
86.0
Visual Cue Quality4% weight
91.0
80.2

Everything else that differs

GPT-5.6 Sol (high)Qwen3.7 Max (default)
Overall / 10088.481.7
Writing Elo23151699
Run-to-run spread (± overall std)1.7503.060
Cost per script (USD)0.1630.054
Avg latency (s)106.4156.8
Open weightsNoNo

Full scorecards: GPT-5.6 Sol (high) · Qwen3.7 Max (default). How scoring works: methodology.

← All comparisons