Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.0 to 84.0.
GPT-5.6 Sol at ultra effort wins this one clearly. It sits at #12 with 2177.4 Elo against DeepSeek V4 Pro's #37 and 1880.7, and the confidence intervals don't come close to touching, so the 296.7-point gap is real. Overall it's 87.97 to 84.04. GPT-5.6 sweeps every metric: the widest gap is visual cues, 88.55 to 79.47, with the numbers-and-slop metric next at 90.13 to 82.09 and continuity and emotion at 85.85 to 80.77. Only hook strength stays close, 87.69 to 86.83. GPT-5.6 is also steadier, with an overall standard deviation of 1.45 against DeepSeek's 4.37, and the records agree: 1042 wins and 10 losses against 674 and 68. What DeepSeek has is the price. It's open weights at about two cents per script against roughly fourteen cents, more than 6x cheaper for a 3.93-point overall gap. If every script gets a human rewrite anyway, that trade is worth considering; if drafts ship close to as-is, it isn't.
Pick GPT-5.6 Sol (ultra) if you want the stronger script on every metric, especially visual cues and consistency from draft to draft, at about fourteen cents each.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights at about two cents per script and can accept a 3.93-point overall gap with noticeably more variance.
Blue bars: GPT-5.6 Sol (ultra). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (ultra) | DeepSeek V4 Pro 0813 (max) | |
|---|---|---|
| Overall / 100 | 88.0 | 84.0 |
| Writing Elo | 2177 | 1881 |
| Run-to-run spread (± overall std) | 1.450 | 4.370 |
| Cost per script (USD) | 0.143 | 0.022 |
| Avg latency (s) | 195.5 | 195.2 |
| Open weights | No | Yes |
Full scorecards: GPT-5.6 Sol (ultra) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.
← All comparisons