Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.4 to 83.4.
Both of these are length specialists, which makes the comparison unusually clean. GPT-5.6 Sol at ultra is #16 with 94.34 length adherence; Qwen3.8 Max is #38 with 92.90. Under a point apart on the metric they both do best. Everything else favours GPT, and the overall is 88.40 against 83.44. The decisive gap is continuity: 85.72 against 78.08, nearly eight points. GPT also takes substance 90.49 to 84.37 and cue quality 91.17 to 85.22. Where Qwen stays competitive is anti-slop at 87.78 against 90.60, so it is not a slop machine, it just does not hold a long script together. Cost is close, 14 cents against 15, and both are slow, 421 seconds against 397. Given near-identical price and speed, and five points of quality, this one is not close on the merits.
Pick GPT-5.6 Sol (ultra). Same price, same speed, five points better, and the flow gap is the one you feel reading it.
Pick Qwen3.8 Max only if you are already on Alibaba infrastructure and want comparable length discipline there.
Blue bars: GPT-5.6 Sol (ultra). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (ultra) | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 88.4 | 83.4 |
| Writing Elo | 2316 | 1866 |
| Run-to-run spread (± overall std) | 1.970 | 5.620 |
| Cost per script (USD) | 0.137 | 0.150 |
| Avg latency (s) | 420.9 | 397.4 |
| Open weights | No | No |
Full scorecards: GPT-5.6 Sol (ultra) · Qwen3.8 Max. How scoring works: methodology.
← All comparisons