Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.5 to 84.1.
GPT-5.6 Sol (high) wins this one cleanly. It sits fourteenth on the board at 2163.5 elo against DeepSeek V4 Pro 0813 (max) at #39 and 1847.7, a gap of 315.8 points, and the confidence intervals do not overlap. The overall scores are closer than the elo suggests, 88.47 to 84.09, but the variance tells the story: GPT-5.6 Sol holds a standard deviation of 1.72 while DeepSeek swings at 6.49. The per-metric spread confirms it. Cue quality is 90.65 against 79.38, and length adherence 91.45 against 81.21. Hook strength is a near tie at 88.11 to 87.99, so DeepSeek can open a script well; it just cannot sustain the structure. The counterweight is price. DeepSeek runs about two cents per script on exact billing, against roughly seventeen cents estimated for GPT-5.6, and its weights are open, so you can host it yourself.
Pick GPT-5.6 Sol (high) if you need consistent, structure-tight scripts where cue quality and length adherence matter more than the roughly seventeen cents each run costs.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights and near-equal hook strength at about two cents per script, and you can tolerate wider run-to-run swings.
Blue bars: GPT-5.6 Sol (high). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | DeepSeek V4 Pro 0813 (max) | |
|---|---|---|
| Overall / 100 | 88.5 | 84.1 |
| Writing Elo | 2164 | 1848 |
| Run-to-run spread (± overall std) | 1.720 | 6.490 |
| Cost per script (USD) | 0.168 | 0.022 |
| Avg latency (s) | 107.3 | 227.3 |
| Open weights | No | Yes |
Full scorecards: GPT-5.6 Sol (high) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.
← All comparisons