Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.5 to 83.3.
GPT-5.6 Sol (high) wins this pairing comfortably. It sits fourteenth on the board at 2163.5 Elo against Qwen3.8 Max at #42 and 1798.7, a gap of 364.8 points, and the confidence intervals do not overlap. The overall scores are closer than the Elo gap suggests: 88.47 against 83.35, a difference of 5.12. Consistency separates them more than the averages do. GPT-5.6 Sol holds a standard deviation of 1.72 while Qwen3.8 Max swings at 5.43. The sharpest per-metric contrast is continuity and emotion, 86.15 against 78.21, and cue quality follows the same pattern at 90.65 against 84.78. Cost barely enters the argument: about seventeen cents per script for GPT-5.6 Sol against about fifteen cents for Qwen3.8 Max, close enough to ignore. Neither model is open weights. The stronger writer here costs roughly the same as the weaker one, which settles it.
Pick GPT-5.6 Sol (high) if you want top-tier consistency, stronger continuity and emotion, and cleaner narration cues for roughly the same money.
Pick Qwen3.8 Max if you want to save about two cents per script and can tolerate wider swings in output quality from run to run.
Blue bars: GPT-5.6 Sol (high). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 88.5 | 83.3 |
| Writing Elo | 2164 | 1799 |
| Run-to-run spread (± overall std) | 1.720 | 5.430 |
| Cost per script (USD) | 0.168 | 0.150 |
| Avg latency (s) | 107.3 | 395.6 |
| Open weights | No | No |
Full scorecards: GPT-5.6 Sol (high) · Qwen3.8 Max. How scoring works: methodology.
← All comparisons