Towards AITowards AIToneBench

MiniMax M3 vs DeepSeek V4 Pro (xhigh)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 80.5.

MiniMax M3
#42 Elo 1790 · 82.4/100
DeepSeek V4 Pro (xhigh)
#60 Elo 1601 · 80.5/100
Cost / script
$0.019 vs $0.014
Human baseline
92.7 MiniMax M3 falls below it · DeepSeek V4 Pro (xhigh) falls below it

The verdict

I cannot call this one cleanly, and I would rather say that than fake a verdict. MiniMax M3 sits ahead on paper, but the confidence intervals overlap, so the honest read is that these two are within noise of each other. What actually separates them is consistency. MiniMax is the swingiest model on these pages: it scored 87.32 on our deployment task and 75.53 on the AI-learning roadmap task, a 12-point spread on the same benchmark. DeepSeek V4 Pro never wowed me but never collapsed either, staying inside a five-point band across the current task set, including the graph engineering script. Price is a wash, both around a cent or two per script, and both are open weights. So this comes down to what you can tolerate. If you generate one script and ship it, DeepSeek's floor protects you. If you generate several candidates and pick the best, MiniMax's ceiling is the higher one. That is genuinely how I would use them.

Pick MiniMax M3 if you generate several candidates per task and keep the best, since its ceiling is clearly higher.
Pick DeepSeek V4 Pro (xhigh) if you run one generation per script and need a predictable floor.

Metric by metric

Blue bars: MiniMax M3. Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.8
82.2
Writing Craft & Clarity13% weight
83.9
82.2
Substance, Accuracy & Value15% weight
81.8
81.0
Continuity & Emotion14% weight
79.2
78.5
YouTube Best Practices12% weight
77.7
76.9
Hook Strength10% weight
86.1
84.6
Length Adherence8% weight
82.6
80.5
Slop Score (EQ-Bench + ours)5% weight
87.3
79.4
Visual Cue Quality4% weight
83.0
73.6

Everything else that differs

MiniMax M3DeepSeek V4 Pro (xhigh)
Overall / 10082.480.5
Writing Elo17901601
Run-to-run spread (± overall std)5.6103.670
Cost per script (USD)0.0190.014
Avg latency (s)185.4163.2
Open weightsYesYes

Full scorecards: MiniMax M3 · DeepSeek V4 Pro (xhigh). How scoring works: methodology.

← All comparisons