Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.5 to 83.3.

GPT-5.6 Sol (high)
#15 Elo 2164 · 88.5/100
Qwen3.8 Max
#41 Elo 1799 · 83.3/100
Cost / script
$0.168 vs $0.150
Human baseline
90.3 GPT-5.6 Sol (high) falls below it · Qwen3.8 Max falls below it

The verdict

GPT-5.6 Sol (high) wins this pairing comfortably. It sits fourteenth on the board at 2163.5 Elo against Qwen3.8 Max at #42 and 1798.7, a gap of 364.8 points, and the confidence intervals do not overlap. The overall scores are closer than the Elo gap suggests: 88.47 against 83.35, a difference of 5.12. Consistency separates them more than the averages do. GPT-5.6 Sol holds a standard deviation of 1.72 while Qwen3.8 Max swings at 5.43. The sharpest per-metric contrast is continuity and emotion, 86.15 against 78.21, and cue quality follows the same pattern at 90.65 against 84.78. Cost barely enters the argument: about seventeen cents per script for GPT-5.6 Sol against about fifteen cents for Qwen3.8 Max, close enough to ignore. Neither model is open weights. The stronger writer here costs roughly the same as the weaker one, which settles it.

Pick GPT-5.6 Sol (high) if you want top-tier consistency, stronger continuity and emotion, and cleaner narration cues for roughly the same money.
Pick Qwen3.8 Max if you want to save about two cents per script and can tolerate wider swings in output quality from run to run.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.5
81.8
Writing Craft & Clarity13% weight
88.3
83.4
Substance, Accuracy & Value15% weight
89.9
84.6
Continuity & Emotion14% weight
86.2
78.2
YouTube Best Practices12% weight
86.7
80.8
Hook Strength10% weight
88.1
84.0
Length Adherence8% weight
91.5
90.6
Slop Score (EQ-Bench + ours)5% weight
93.4
92.0
Visual Cue Quality4% weight
90.7
84.8

Everything else that differs

GPT-5.6 Sol (high)Qwen3.8 Max
Overall / 10088.583.3
Writing Elo21641799
Run-to-run spread (± overall std)1.7205.430
Cost per script (USD)0.1680.150
Avg latency (s)107.3395.6
Open weightsNoNo

Full scorecards: GPT-5.6 Sol (high) · Qwen3.8 Max. How scoring works: methodology.

← All comparisons