Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs DeepSeek V4 Pro (xhigh)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 80.5.

GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
DeepSeek V4 Pro (xhigh)
#60 Elo 1601 · 80.5/100
Cost / script
$0.163 vs $0.014
Human baseline
92.7 GPT-5.6 Sol (high) falls below it · DeepSeek V4 Pro (xhigh) falls below it

The verdict

GPT-5.6 Sol at high effort wins comfortably, and I owe you a disclosure before the numbers: Sol also sits on my three-judge panel. Claude Opus 5 and DeepSeek V4 Flash score independently though, and the consensus matches what I see reading the scripts myself. The Elo gap clears both confidence intervals, so no hedging needed. The metric that decides it for me is visual cues, 91.02 against 73.61. Sol writes [SHOW:] cues I could hand straight to my editor; DeepSeek's cues are vague enough that I would rewrite most of them. It is also the faster model here, returning a script in about a minute and a half while DeepSeek takes closer to three. Cost is where DeepSeek fights back, about a cent and a half per script versus 16 cents, plus open weights. That is a real argument for high-volume drafting. For scripts a human barely touches, Sol is the safer bet.

Pick GPT-5.6 Sol (high) if you want near-top scripts with visual cues your editor can use as-is, at 16 cents a run.
Pick DeepSeek V4 Pro (xhigh) if cost rules the decision and you are fine rewriting the [SHOW:] cues yourself.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.3
82.2
Writing Craft & Clarity13% weight
88.1
82.2
Substance, Accuracy & Value15% weight
90.1
81.0
Continuity & Emotion14% weight
86.1
78.5
YouTube Best Practices12% weight
86.5
76.9
Hook Strength10% weight
87.5
84.6
Length Adherence8% weight
91.7
80.5
Slop Score (EQ-Bench + ours)5% weight
93.4
79.4
Visual Cue Quality4% weight
91.0
73.6

Everything else that differs

GPT-5.6 Sol (high)DeepSeek V4 Pro (xhigh)
Overall / 10088.480.5
Writing Elo23151601
Run-to-run spread (± overall std)1.7503.670
Cost per script (USD)0.1630.014
Avg latency (s)106.4163.2
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (high) · DeepSeek V4 Pro (xhigh). How scoring works: methodology.

← All comparisons