Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs Qwen3.7 Max (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 81.5.

GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
Cost / script
$0.163 vs $0.054
Human baseline
92.7 GPT-5.6 Sol (high) falls below it · Qwen3.7 Max (high) falls below it

The verdict

This one isn't close. GPT-5.6 wins by more than five points overall and the Elo intervals sit miles apart. Qwen3.7 Max opens well, its hook score of 85.0 is respectable, but everything after the hook slides: writing craft, substance, and emotional continuity all land in the low 80s or worse, while GPT stays mid 80s to low 90s. The ugliest gap is visual cues, 81.0 against 91.0, so Qwen's on-screen directions need real cleanup before an editor can use them. What makes this an easy call for me is that Qwen doesn't buy its way out either. It's a closed model, same as GPT, and its estimated five cents per script is only about a third of GPT's fifteen. Cheaper, sure, but not open and not cheap enough to justify the quality drop. And I'll say the fair thing: strong hooks are a real skill, and Qwen has them. But a strong hook on a mediocre body is exactly the script that gets clicked and then abandoned. GPT-5.6, easily.

Pick GPT-5.6 Sol if you care about everything after the first thirty seconds; the price gap is worth it.
Pick Qwen3.7 Max if you mostly need hooks and intros at a third of the cost and will rewrite the body anyway.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.3
80.9
Writing Craft & Clarity13% weight
88.1
81.5
Substance, Accuracy & Value15% weight
90.1
81.2
Continuity & Emotion14% weight
86.1
79.2
YouTube Best Practices12% weight
86.5
80.7
Hook Strength10% weight
87.5
85.0
Length Adherence8% weight
91.7
83.3
Slop Score (EQ-Bench + ours)5% weight
93.4
85.9
Visual Cue Quality4% weight
91.0
78.4

Everything else that differs

GPT-5.6 Sol (high)Qwen3.7 Max (high)
Overall / 10088.481.5
Writing Elo23151674
Run-to-run spread (± overall std)1.7502.580
Cost per script (USD)0.1630.054
Avg latency (s)106.4157.2
Open weightsNoNo

Full scorecards: GPT-5.6 Sol (high) · Qwen3.7 Max (high). How scoring works: methodology.

← All comparisons