Towards AITowards AIToneBench

GPT-5.6 Sol (ultra) vs DeepSeek V4 Pro 0813 (max)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.0 to 84.0.

GPT-5.6 Sol (ultra)
#12 Elo 2177 · 88.0/100
DeepSeek V4 Pro 0813 (max)
#37 Elo 1881 · 84.0/100
Cost / script
$0.143 vs $0.022
Human baseline
90.2 GPT-5.6 Sol (ultra) falls below it · DeepSeek V4 Pro 0813 (max) falls below it

The verdict

GPT-5.6 Sol at ultra effort wins this one clearly. It sits at #12 with 2177.4 Elo against DeepSeek V4 Pro's #37 and 1880.7, and the confidence intervals don't come close to touching, so the 296.7-point gap is real. Overall it's 87.97 to 84.04. GPT-5.6 sweeps every metric: the widest gap is visual cues, 88.55 to 79.47, with the numbers-and-slop metric next at 90.13 to 82.09 and continuity and emotion at 85.85 to 80.77. Only hook strength stays close, 87.69 to 86.83. GPT-5.6 is also steadier, with an overall standard deviation of 1.45 against DeepSeek's 4.37, and the records agree: 1042 wins and 10 losses against 674 and 68. What DeepSeek has is the price. It's open weights at about two cents per script against roughly fourteen cents, more than 6x cheaper for a 3.93-point overall gap. If every script gets a human rewrite anyway, that trade is worth considering; if drafts ship close to as-is, it isn't.

Pick GPT-5.6 Sol (ultra) if you want the stronger script on every metric, especially visual cues and consistency from draft to draft, at about fourteen cents each.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights at about two cents per script and can accept a 3.93-point overall gap with noticeably more variance.

Metric by metric

Blue bars: GPT-5.6 Sol (ultra). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
86.8
84.7
Writing Craft & Clarity13% weight
87.7
84.9
Substance, Accuracy & Value15% weight
89.0
84.7
Continuity & Emotion14% weight
85.8
80.8
YouTube Best Practices12% weight
86.5
81.1
Hook Strength10% weight
87.7
86.8
Length Adherence8% weight
91.9
86.2
Slop Score (EQ-Bench + ours)5% weight
93.3
88.0
Visual Cue Quality4% weight
88.5
79.5

Everything else that differs

GPT-5.6 Sol (ultra)DeepSeek V4 Pro 0813 (max)
Overall / 10088.084.0
Writing Elo21771881
Run-to-run spread (± overall std)1.4504.370
Cost per script (USD)0.1430.022
Avg latency (s)195.5195.2
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (ultra) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.

← All comparisons