Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 80.5.
GPT-5.6 Sol at high effort wins comfortably, and I owe you a disclosure before the numbers: Sol also sits on my three-judge panel. Claude Opus 5 and DeepSeek V4 Flash score independently though, and the consensus matches what I see reading the scripts myself. The Elo gap clears both confidence intervals, so no hedging needed. The metric that decides it for me is visual cues, 91.02 against 73.61. Sol writes [SHOW:] cues I could hand straight to my editor; DeepSeek's cues are vague enough that I would rewrite most of them. It is also the faster model here, returning a script in about a minute and a half while DeepSeek takes closer to three. Cost is where DeepSeek fights back, about a cent and a half per script versus 16 cents, plus open weights. That is a real argument for high-volume drafting. For scripts a human barely touches, Sol is the safer bet.
Pick GPT-5.6 Sol (high) if you want near-top scripts with visual cues your editor can use as-is, at 16 cents a run.
Pick DeepSeek V4 Pro (xhigh) if cost rules the decision and you are fine rewriting the [SHOW:] cues yourself.
Blue bars: GPT-5.6 Sol (high). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | DeepSeek V4 Pro (xhigh) | |
|---|---|---|
| Overall / 100 | 88.4 | 80.5 |
| Writing Elo | 2315 | 1601 |
| Run-to-run spread (± overall std) | 1.750 | 3.670 |
| Cost per script (USD) | 0.163 | 0.014 |
| Avg latency (s) | 106.4 | 163.2 |
| Open weights | No | Yes |
Full scorecards: GPT-5.6 Sol (high) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.
← All comparisons