Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs DeepSeek V4 Pro 0813 (max)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.5 to 84.1.

GPT-5.6 Sol (high)
#15 Elo 2164 · 88.5/100
DeepSeek V4 Pro 0813 (max)
#38 Elo 1848 · 84.1/100
Cost / script
$0.168 vs $0.022
Human baseline
90.3 GPT-5.6 Sol (high) falls below it · DeepSeek V4 Pro 0813 (max) falls below it

The verdict

GPT-5.6 Sol (high) wins this one cleanly. It sits fourteenth on the board at 2163.5 elo against DeepSeek V4 Pro 0813 (max) at #39 and 1847.7, a gap of 315.8 points, and the confidence intervals do not overlap. The overall scores are closer than the elo suggests, 88.47 to 84.09, but the variance tells the story: GPT-5.6 Sol holds a standard deviation of 1.72 while DeepSeek swings at 6.49. The per-metric spread confirms it. Cue quality is 90.65 against 79.38, and length adherence 91.45 against 81.21. Hook strength is a near tie at 88.11 to 87.99, so DeepSeek can open a script well; it just cannot sustain the structure. The counterweight is price. DeepSeek runs about two cents per script on exact billing, against roughly seventeen cents estimated for GPT-5.6, and its weights are open, so you can host it yourself.

Pick GPT-5.6 Sol (high) if you need consistent, structure-tight scripts where cue quality and length adherence matter more than the roughly seventeen cents each run costs.
Pick DeepSeek V4 Pro 0813 (max) if you want open weights and near-equal hook strength at about two cents per script, and you can tolerate wider run-to-run swings.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: DeepSeek V4 Pro 0813 (max). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.5
85.6
Writing Craft & Clarity13% weight
88.3
85.6
Substance, Accuracy & Value15% weight
89.9
83.4
Continuity & Emotion14% weight
86.2
81.8
YouTube Best Practices12% weight
86.7
81.9
Hook Strength10% weight
88.1
88.0
Length Adherence8% weight
91.5
81.2
Slop Score (EQ-Bench + ours)5% weight
93.4
88.8
Visual Cue Quality4% weight
90.7
79.4

Everything else that differs

GPT-5.6 Sol (high)DeepSeek V4 Pro 0813 (max)
Overall / 10088.584.1
Writing Elo21641848
Run-to-run spread (± overall std)1.7206.490
Cost per script (USD)0.1680.022
Avg latency (s)107.3227.3
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (high) · DeepSeek V4 Pro 0813 (max). How scoring works: methodology.

← All comparisons