Towards AITowards AIToneBench

GPT-5.6 Sol (ultra) vs DeepSeek V4 Flash 0731

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.4 to 84.2.

GPT-5.6 Sol (ultra)
#16 Elo 2316 · 88.4/100
DeepSeek V4 Flash 0731
#34 Elo 1955 · 84.2/100
Cost / script
$0.137 vs $0.005
Human baseline
92.7 GPT-5.6 Sol (ultra) falls below it · DeepSeek V4 Flash 0731 falls below it

The verdict

GPT-5.6 Sol at ultra is #16 and DeepSeek V4 Flash 0731 is #34, so 88.40 against 84.23. Five and a half points. GPT wins almost everywhere, and widest on the mechanical metrics: length adherence 94.34 against 83.57, fourteen points, and cue quality 91.17 against 82.40. Substance too, 90.49 against 83.48. DeepSeek's answer is hooks, where it is genuinely strong at 87.55 against GPT's 86.39, and voice at 85.44 is closer than the rank gap implies. Then the economics, which are the real story here: half a cent a script against 14 cents. GPT is twenty eight times more expensive, and DeepSeek is open weights. The 0731 build also jumped 45 places over the 0423 one, so this is a fast-moving model, not a static budget option.

Pick GPT-5.6 Sol (ultra) when length and cue precision matter and the cost per script is noise.
Pick DeepSeek V4 Flash 0731 when volume matters. Twenty eight times cheaper, open, with the better hooks of the two.

Metric by metric

Blue bars: GPT-5.6 Sol (ultra). Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.0
85.4
Writing Craft & Clarity13% weight
88.1
85.3
Substance, Accuracy & Value15% weight
90.5
83.5
Continuity & Emotion14% weight
85.7
82.0
YouTube Best Practices12% weight
86.1
81.1
Hook Strength10% weight
86.4
87.5
Length Adherence8% weight
94.3
83.6
Slop Score (EQ-Bench + ours)5% weight
93.5
88.8
Visual Cue Quality4% weight
91.2
82.4

Everything else that differs

GPT-5.6 Sol (ultra)DeepSeek V4 Flash 0731
Overall / 10088.484.2
Writing Elo23161955
Run-to-run spread (± overall std)1.9709.240
Cost per script (USD)0.1370.005
Avg latency (s)420.9289.0
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (ultra) · DeepSeek V4 Flash 0731. How scoring works: methodology.

← All comparisons