Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.4 to 84.2.
GPT-5.6 Sol at ultra is #16 and DeepSeek V4 Flash 0731 is #34, so 88.40 against 84.23. Five and a half points. GPT wins almost everywhere, and widest on the mechanical metrics: length adherence 94.34 against 83.57, fourteen points, and cue quality 91.17 against 82.40. Substance too, 90.49 against 83.48. DeepSeek's answer is hooks, where it is genuinely strong at 87.55 against GPT's 86.39, and voice at 85.44 is closer than the rank gap implies. Then the economics, which are the real story here: half a cent a script against 14 cents. GPT is twenty eight times more expensive, and DeepSeek is open weights. The 0731 build also jumped 45 places over the 0423 one, so this is a fast-moving model, not a static budget option.
Pick GPT-5.6 Sol (ultra) when length and cue precision matter and the cost per script is noise.
Pick DeepSeek V4 Flash 0731 when volume matters. Twenty eight times cheaper, open, with the better hooks of the two.
Blue bars: GPT-5.6 Sol (ultra). Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (ultra) | DeepSeek V4 Flash 0731 | |
|---|---|---|
| Overall / 100 | 88.4 | 84.2 |
| Writing Elo | 2316 | 1955 |
| Run-to-run spread (± overall std) | 1.970 | 9.240 |
| Cost per script (USD) | 0.137 | 0.005 |
| Avg latency (s) | 420.9 | 289.0 |
| Open weights | No | Yes |
Full scorecards: GPT-5.6 Sol (ultra) · DeepSeek V4 Flash 0731. How scoring works: methodology.
← All comparisons