Towards AITowards AIToneBench

GPT-5.6 Sol (ultra) vs DeepSeek V4.1 Flash (max)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.0 to 86.8. Their confidence intervals overlap, so treat the order as close rather than settled.

GPT-5.6 Sol (ultra)
#16 Elo 2145 · 88.0/100
DeepSeek V4.1 Flash (max)
#22 Elo 2085 · 86.8/100
Cost / script
$0.363 vs $0.012
Human baseline
90.2 GPT-5.6 Sol (ultra) falls below it · DeepSeek V4.1 Flash (max) falls below it

The verdict

GPT-5.6 Sol (ultra) has the stronger current result: #16 at 2145.0 Elo and 87.97 overall, compared with DeepSeek V4.1 Flash (max) at #22, 2085.1 Elo, and 86.75 overall. Their 95% Elo confidence intervals overlap, so the exact ordering should be treated as uncertain (2104.90–2175.90 and 2037.20–2122.70). GPT-5.6 Sol (ultra) has its clearest metric edges in Visual Cue Quality (88.55 versus 82.35) and Substance, Accuracy & Value (89.01 versus 85.23). DeepSeek V4.1 Flash (max) counters on Hook Strength (88.88 versus 87.69) and Tone & Voice Match (87.42 versus 86.80). At the measured run mix, GPT-5.6 Sol (ultra) costs $0.363 per article versus $0.012 for DeepSeek V4.1 Flash (max); DeepSeek V4.1 Flash (max) is the cheaper route. DeepSeek V4.1 Flash (max) is the open-weights option; GPT-5.6 Sol (ultra) is closed. On the current automated evidence, GPT-5.6 Sol (ultra) is the stronger default; DeepSeek V4.1 Flash (max) remains a defensible choice when its specific strengths, price, or deployment profile matter more than the headline rank.

Pick GPT-5.6 Sol (ultra) when you prioritize the stronger current board result, visual cue quality, and substance, accuracy & value.
Pick DeepSeek V4.1 Flash (max) when you prioritize hook strength, tone & voice match, open weights, and lower measured cost.

Metric by metric

Blue bars: GPT-5.6 Sol (ultra). Orange bars: DeepSeek V4.1 Flash (max). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
86.8
87.4
Writing Craft & Clarity13% weight
87.7
87.3
Substance, Accuracy & Value15% weight
89.0
85.2
Continuity & Emotion14% weight
85.8
84.9
YouTube Best Practices12% weight
86.5
84.9
Hook Strength10% weight
87.7
88.9
Length Adherence8% weight
91.9
90.0
Slop Score (EQ-Bench + ours)5% weight
93.3
91.0
Visual Cue Quality4% weight
88.5
82.3

Everything else that differs

GPT-5.6 Sol (ultra)DeepSeek V4.1 Flash (max)
Overall / 10088.086.8
Writing Elo21452085
Run-to-run spread (± overall std)1.4506.940
Cost per script (USD)0.3630.012
Avg latency (s)195.574.6
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (ultra) · DeepSeek V4.1 Flash (max). How scoring works: methodology.

← All comparisons