Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 80.5.
I cannot call this one cleanly, and I would rather say that than fake a verdict. MiniMax M3 sits ahead on paper, but the confidence intervals overlap, so the honest read is that these two are within noise of each other. What actually separates them is consistency. MiniMax is the swingiest model on these pages: it scored 87.32 on our deployment task and 75.53 on the AI-learning roadmap task, a 12-point spread on the same benchmark. DeepSeek V4 Pro never wowed me but never collapsed either, staying inside a five-point band across the current task set, including the graph engineering script. Price is a wash, both around a cent or two per script, and both are open weights. So this comes down to what you can tolerate. If you generate one script and ship it, DeepSeek's floor protects you. If you generate several candidates and pick the best, MiniMax's ceiling is the higher one. That is genuinely how I would use them.
Pick MiniMax M3 if you generate several candidates per task and keep the best, since its ceiling is clearly higher.
Pick DeepSeek V4 Pro (xhigh) if you run one generation per script and need a predictable floor.
Blue bars: MiniMax M3. Orange bars: DeepSeek V4 Pro (xhigh). Same 0–100 scale; the bold bar wins that metric.
| MiniMax M3 | DeepSeek V4 Pro (xhigh) | |
|---|---|---|
| Overall / 100 | 82.4 | 80.5 |
| Writing Elo | 1790 | 1601 |
| Run-to-run spread (± overall std) | 5.610 | 3.670 |
| Cost per script (USD) | 0.019 | 0.014 |
| Avg latency (s) | 185.4 | 163.2 |
| Open weights | Yes | Yes |
Full scorecards: MiniMax M3 · DeepSeek V4 Pro (xhigh). How scoring works: methodology.
← All comparisons