Towards AITowards AIToneBench

Grok 4.6 (high) vs DeepSeek V4 Flash 0731

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 83.0.

Grok 4.6 (high)
#18 Elo 2144 · 88.1/100
DeepSeek V4 Flash 0731
#40 Elo 1806 · 83.0/100
Cost / script
$0.115 vs $0.005
Human baseline
90.3 Grok 4.6 (high) falls below it · DeepSeek V4 Flash 0731 falls below it

The verdict

Grok 4.6 wins this one comfortably. Seventeenth on the board against #45, Elo 2144.5 to 1806.5, and the confidence intervals sit far apart, so the ranking is real. The overall scores tell the same story: 88.14 against 83.05, a 5.09-point gap, and DeepSeek's 9.42 standard deviation means its scripts swing wildly where Grok's 1.43 stays steady. The biggest per-metric gaps are structural: YouTube best practices, 87.08 to 79.86, and numbers handling, 88.24 to 80.85. DeepSeek does take one metric outright, slop scoring, 95.68 to Grok's 95.44, a whisker but a genuine win. Then the ledger. DeepSeek costs about half a cent per script against Grok's about twelve cents, both exact figures, roughly 24x cheaper, and it is open weights. That buys a lot of forgiveness for a 5.09-point gap if you draft at volume and edit anyway.

Pick Grok 4.6 (high) if you want the steadier, higher-scoring script on every run and about twelve cents per draft is a fair price for skipping heavy edits.
Pick DeepSeek V4 Flash 0731 if you draft at volume with an editor in the loop and want open weights at about half a cent per script.

Metric by metric

Blue bars: Grok 4.6 (high). Orange bars: DeepSeek V4 Flash 0731. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
88.8
84.2
Writing Craft & Clarity13% weight
88.0
84.2
Substance, Accuracy & Value15% weight
88.0
82.4
Continuity & Emotion14% weight
86.3
80.5
YouTube Best Practices12% weight
87.1
79.9
Hook Strength10% weight
90.0
87.0
Length Adherence8% weight
88.3
81.6
Slop Score (EQ-Bench + ours)5% weight
91.8
88.3
Visual Cue Quality4% weight
86.4
81.0

Everything else that differs

Grok 4.6 (high)DeepSeek V4 Flash 0731
Overall / 10088.183.0
Writing Elo21441806
Run-to-run spread (± overall std)1.4309.420
Cost per script (USD)0.1150.005
Avg latency (s)199.2276.4
Open weightsNoYes

Full scorecards: Grok 4.6 (high) · DeepSeek V4 Flash 0731. How scoring works: methodology.

← All comparisons