Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Pro 0813 (max) leads overall, 84.1 to 82.2. Their confidence intervals overlap, so treat the order as close rather than settled.
DeepSeek V4 Pro 0813 (max) takes this one. It sits at #39 with an elo of 1847.7 against MiniMax M3 at #51 and 1734.6, a gap of 113.1 points, and the confidence intervals do not overlap. The overalls are closer than the elo suggests: 84.09 against 82.19, a difference of 1.90. The separation comes from delivery. DeepSeek leads continuity and emotion 81.85 to 78.50 and anti-slop numbers 83.59 to 78.24. MiniMax answers with stronger cue quality, 82.62 to 79.38, and tighter length adherence at 83.11. Cost settles nothing here: both run about two cents per script. Both ship open weights, so either can be self-hosted. A clear win on the aggregate, with real trade-offs underneath.
Pick DeepSeek V4 Pro 0813 (max) if you want the stronger aggregate writer, with better hooks, continuity, and number handling at effectively the same roughly two-cent price.
Pick MiniMax M3 if cue quality and length adherence matter most to your workflow and you will trade a lower overall score for them at near-identical cost.
Blue bars: DeepSeek V4 Pro 0813 (max). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| DeepSeek V4 Pro 0813 (max) | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 84.1 | 82.2 |
| Writing Elo | 1848 | 1735 |
| Run-to-run spread (± overall std) | 6.490 | 5.220 |
| Cost per script (USD) | 0.022 | 0.020 |
| Avg latency (s) | 227.3 | 201.6 |
| Open weights | Yes | Yes |
Full scorecards: DeepSeek V4 Pro 0813 (max) · MiniMax M3. How scoring works: methodology.
← All comparisons