Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 81.7.
GPT-5.6 Sol at high reasoning is #17, Qwen3.7 Max is #50, and the overall is 88.38 against 81.73. Just under seven points. The gap is widest on the mechanical metrics, which is GPT's whole personality on this board: anti-slop and numbers 90.60 against 77.84, cue quality 91.02 against 80.24. Both over twelve points. If your scripts are full of figures and visual direction, that's the entire argument. Continuity is the third big one, 86.07 to 79.03. Qwen's best number is length adherence at 85.79, and it's genuinely close to GPT's 91.69 by the standards of the rest of this comparison. On price and speed GPT is 16 cents and 106 seconds against Qwen's five cents and 157 seconds. So GPT is three times the price and noticeably faster, which is an unusual combination and worth knowing.
Pick GPT-5.6 Sol (high) if numbers and visual cues have to be right, and you want drafts back fast.
Pick Qwen3.7 Max if cost is the driver and you can absorb the slop and cue cleanup.
Blue bars: GPT-5.6 Sol (high). Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | Qwen3.7 Max (default) | |
|---|---|---|
| Overall / 100 | 88.4 | 81.7 |
| Writing Elo | 2315 | 1699 |
| Run-to-run spread (± overall std) | 1.750 | 3.060 |
| Cost per script (USD) | 0.163 | 0.054 |
| Avg latency (s) | 106.4 | 156.8 |
| Open weights | No | No |
Full scorecards: GPT-5.6 Sol (high) · Qwen3.7 Max (default). How scoring works: methodology.
← All comparisons