Towards AITowards AIToneBench

DeepSeek V4.1 Flash (max) vs GPT-5.6 Sol (ultra)

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4.1 Flash (max) leads overall, 88.0 to 88.0. Their confidence intervals overlap, so treat the order as close rather than settled.

DeepSeek V4.1 Flash (max)
#14 Elo 2227 · 88.0/100
GPT-5.6 Sol (ultra)
#19 Elo 2218 · 88.0/100
Cost / script
$0.013 vs $0.363
Human baseline
90.2 DeepSeek V4.1 Flash (max) falls below it · GPT-5.6 Sol (ultra) falls below it

The verdict

This one is closer than the ranks suggest. DeepSeek V4.1 Flash (max) sits at #14 with 2226.8 Elo and 88.03 overall; GPT-5.6 Sol (ultra) is at #19 with 2218.2 Elo and 87.97 overall. Their 95% Elo intervals overlap (2185.5 to 2260.8 against 2177.5 to 2256.6), so do not read much into the exact order. Where DeepSeek V4.1 Flash (max) pulls ahead: Hook Strength (89.07 vs 87.69) and Tone & Voice Match (88.06 vs 86.80). GPT-5.6 Sol (ultra) still wins on Visual Cue Quality (88.55 vs 83.68) and Slop Score (EQ-Bench + ours) (93.34 vs 91.36), so it is not a clean sweep. Price points the same way: DeepSeek V4.1 Flash (max) costs about $0.013 per article against $0.363 for GPT-5.6 Sol (ultra). DeepSeek V4.1 Flash (max) publishes open weights; GPT-5.6 Sol (ultra) does not. Either is a reasonable default: lean DeepSeek V4.1 Flash (max) for hook strength, voice match, the lower price, and open weights, GPT-5.6 Sol (ultra) for visual cues and slop score.

Pick GPT-5.6 Sol (ultra) for visual cues and slop score.
Pick DeepSeek V4.1 Flash (max) for the stronger board result, hook strength, voice match, open weights, and the lower price.

Metric by metric

Blue bars: DeepSeek V4.1 Flash (max). Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
88.1
86.8
Writing Craft & Clarity13% weight
88.4
87.7
Substance, Accuracy & Value15% weight
87.2
89.0
Continuity & Emotion14% weight
86.6
85.8
YouTube Best Practices12% weight
86.9
86.5
Hook Strength10% weight
89.1
87.7
Length Adherence8% weight
92.0
91.9
Slop Score (EQ-Bench + ours)5% weight
91.4
93.3
Visual Cue Quality4% weight
83.7
88.5

Everything else that differs

DeepSeek V4.1 Flash (max)GPT-5.6 Sol (ultra)
Overall / 10088.088.0
Writing Elo22272218
Run-to-run spread (± overall std)1.2801.450
Cost per script (USD)0.0130.363
Avg latency (s)76.2195.5
Open weightsYesNo

Full scorecards: DeepSeek V4.1 Flash (max) · GPT-5.6 Sol (ultra). How scoring works: methodology.

← All comparisons