Towards AITowards AIToneBench

DeepSeek V4.1 Flash (max) vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4.1 Flash (max) leads overall, 86.8 to 83.4.

DeepSeek V4.1 Flash (max)
#22 Elo 2085 · 86.8/100
MiniMax M3
#48 Elo 1760 · 83.4/100
Cost / script
$0.012 vs $0.023
Human baseline
90.2 DeepSeek V4.1 Flash (max) falls below it · MiniMax M3 falls below it

The verdict

DeepSeek V4.1 Flash (max) has the stronger current result: #22 at 2085.1 Elo and 86.75 overall, compared with MiniMax M3 at #48, 1760.5 Elo, and 83.44 overall. Their 95% Elo confidence intervals do not overlap, which supports the directional ordering (2037.20–2122.70 and 1716.10–1801.80). DeepSeek V4.1 Flash (max) has its clearest metric edges in Slop Score (EQ-Bench + ours) (91.02 versus 86.04) and YouTube Best Practices (84.92 versus 80.43). MiniMax M3 does not lead an individual published metric in this pairing. At the measured run mix, DeepSeek V4.1 Flash (max) costs $0.012 per article versus $0.023 for MiniMax M3; DeepSeek V4.1 Flash (max) is the cheaper route. Both models publish open weights. On the current automated evidence, DeepSeek V4.1 Flash (max) is the stronger default; MiniMax M3 remains a defensible choice when its specific strengths, price, or deployment profile matter more than the headline rank.

Pick DeepSeek V4.1 Flash (max) when you prioritize the stronger current board result, slop score (eq-bench + ours), youtube best practices, open weights, and lower measured cost.
Pick MiniMax M3 when you prioritize open weights.

Metric by metric

Blue bars: DeepSeek V4.1 Flash (max). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.4
83.3
Writing Craft & Clarity13% weight
87.3
84.4
Substance, Accuracy & Value15% weight
85.2
82.2
Continuity & Emotion14% weight
84.9
80.5
YouTube Best Practices12% weight
84.9
80.4
Hook Strength10% weight
88.9
86.5
Length Adherence8% weight
90.0
89.7
Slop Score (EQ-Bench + ours)5% weight
91.0
86.0
Visual Cue Quality4% weight
82.3
81.2

Everything else that differs

DeepSeek V4.1 Flash (max)MiniMax M3
Overall / 10086.883.4
Writing Elo20851760
Run-to-run spread (± overall std)6.9402.030
Cost per script (USD)0.0120.023
Avg latency (s)74.6109.9
Open weightsYesYes

Full scorecards: DeepSeek V4.1 Flash (max) · MiniMax M3. How scoring works: methodology.

← All comparisons