Towards AITowards AIToneBench

DeepSeek V4 Pro 0813 (max) vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Pro 0813 (max) leads overall, 84.1 to 82.2. Their confidence intervals overlap, so treat the order as close rather than settled.

DeepSeek V4 Pro 0813 (max)
#38 Elo 1848 · 84.1/100
MiniMax M3
#46 Elo 1735 · 82.2/100
Cost / script
$0.022 vs $0.020
Human baseline
90.3 DeepSeek V4 Pro 0813 (max) falls below it · MiniMax M3 falls below it

The verdict

DeepSeek V4 Pro 0813 (max) takes this one. It sits at #39 with an elo of 1847.7 against MiniMax M3 at #51 and 1734.6, a gap of 113.1 points, and the confidence intervals do not overlap. The overalls are closer than the elo suggests: 84.09 against 82.19, a difference of 1.90. The separation comes from delivery. DeepSeek leads continuity and emotion 81.85 to 78.50 and anti-slop numbers 83.59 to 78.24. MiniMax answers with stronger cue quality, 82.62 to 79.38, and tighter length adherence at 83.11. Cost settles nothing here: both run about two cents per script. Both ship open weights, so either can be self-hosted. A clear win on the aggregate, with real trade-offs underneath.

Pick DeepSeek V4 Pro 0813 (max) if you want the stronger aggregate writer, with better hooks, continuity, and number handling at effectively the same roughly two-cent price.
Pick MiniMax M3 if cue quality and length adherence matter most to your workflow and you will trade a lower overall score for them at near-identical cost.

Metric by metric

Blue bars: DeepSeek V4 Pro 0813 (max). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
85.6
83.4
Writing Craft & Clarity13% weight
85.6
83.5
Substance, Accuracy & Value15% weight
83.4
81.6
Continuity & Emotion14% weight
81.8
78.5
YouTube Best Practices12% weight
81.9
78.0
Hook Strength10% weight
88.0
86.0
Length Adherence8% weight
81.2
83.1
Slop Score (EQ-Bench + ours)5% weight
88.8
86.9
Visual Cue Quality4% weight
79.4
82.6

Everything else that differs

DeepSeek V4 Pro 0813 (max)MiniMax M3
Overall / 10084.182.2
Writing Elo18481735
Run-to-run spread (± overall std)6.4905.220
Cost per script (USD)0.0220.020
Avg latency (s)227.3201.6
Open weightsYesYes

Full scorecards: DeepSeek V4 Pro 0813 (max) · MiniMax M3. How scoring works: methodology.

← All comparisons