Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 81.5.
This one isn't close. GPT-5.6 wins by more than five points overall and the Elo intervals sit miles apart. Qwen3.7 Max opens well, its hook score of 85.0 is respectable, but everything after the hook slides: writing craft, substance, and emotional continuity all land in the low 80s or worse, while GPT stays mid 80s to low 90s. The ugliest gap is visual cues, 81.0 against 91.0, so Qwen's on-screen directions need real cleanup before an editor can use them. What makes this an easy call for me is that Qwen doesn't buy its way out either. It's a closed model, same as GPT, and its estimated five cents per script is only about a third of GPT's fifteen. Cheaper, sure, but not open and not cheap enough to justify the quality drop. And I'll say the fair thing: strong hooks are a real skill, and Qwen has them. But a strong hook on a mediocre body is exactly the script that gets clicked and then abandoned. GPT-5.6, easily.
Pick GPT-5.6 Sol if you care about everything after the first thirty seconds; the price gap is worth it.
Pick Qwen3.7 Max if you mostly need hooks and intros at a third of the cost and will rewrite the body anyway.
Blue bars: GPT-5.6 Sol (high). Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | Qwen3.7 Max (high) | |
|---|---|---|
| Overall / 100 | 88.4 | 81.5 |
| Writing Elo | 2315 | 1674 |
| Run-to-run spread (± overall std) | 1.750 | 2.580 |
| Cost per script (USD) | 0.163 | 0.054 |
| Avg latency (s) | 106.4 | 157.2 |
| Open weights | No | No |
Full scorecards: GPT-5.6 Sol (high) · Qwen3.7 Max (high). How scoring works: methodology.
← All comparisons