Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4.1 Flash (max) leads overall, 88.0 to 88.0. Their confidence intervals overlap, so treat the order as close rather than settled.
This one is closer than the ranks suggest. DeepSeek V4.1 Flash (max) sits at #14 with 2226.8 Elo and 88.03 overall; GPT-5.6 Sol (ultra) is at #19 with 2218.2 Elo and 87.97 overall. Their 95% Elo intervals overlap (2185.5 to 2260.8 against 2177.5 to 2256.6), so do not read much into the exact order. Where DeepSeek V4.1 Flash (max) pulls ahead: Hook Strength (89.07 vs 87.69) and Tone & Voice Match (88.06 vs 86.80). GPT-5.6 Sol (ultra) still wins on Visual Cue Quality (88.55 vs 83.68) and Slop Score (EQ-Bench + ours) (93.34 vs 91.36), so it is not a clean sweep. Price points the same way: DeepSeek V4.1 Flash (max) costs about $0.013 per article against $0.363 for GPT-5.6 Sol (ultra). DeepSeek V4.1 Flash (max) publishes open weights; GPT-5.6 Sol (ultra) does not. Either is a reasonable default: lean DeepSeek V4.1 Flash (max) for hook strength, voice match, the lower price, and open weights, GPT-5.6 Sol (ultra) for visual cues and slop score.
Pick GPT-5.6 Sol (ultra) for visual cues and slop score.
Pick DeepSeek V4.1 Flash (max) for the stronger board result, hook strength, voice match, open weights, and the lower price.
Blue bars: DeepSeek V4.1 Flash (max). Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.
| DeepSeek V4.1 Flash (max) | GPT-5.6 Sol (ultra) | |
|---|---|---|
| Overall / 100 | 88.0 | 88.0 |
| Writing Elo | 2227 | 2218 |
| Run-to-run spread (± overall std) | 1.280 | 1.450 |
| Cost per script (USD) | 0.013 | 0.363 |
| Avg latency (s) | 76.2 | 195.5 |
| Open weights | Yes | No |
Full scorecards: DeepSeek V4.1 Flash (max) · GPT-5.6 Sol (ultra). How scoring works: methodology.
← All comparisons