Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4.1 Flash (max) leads overall, 86.8 to 83.4.
DeepSeek V4.1 Flash (max) has the stronger current result: #22 at 2085.1 Elo and 86.75 overall, compared with MiniMax M3 at #48, 1760.5 Elo, and 83.44 overall. Their 95% Elo confidence intervals do not overlap, which supports the directional ordering (2037.20–2122.70 and 1716.10–1801.80). DeepSeek V4.1 Flash (max) has its clearest metric edges in Slop Score (EQ-Bench + ours) (91.02 versus 86.04) and YouTube Best Practices (84.92 versus 80.43). MiniMax M3 does not lead an individual published metric in this pairing. At the measured run mix, DeepSeek V4.1 Flash (max) costs $0.012 per article versus $0.023 for MiniMax M3; DeepSeek V4.1 Flash (max) is the cheaper route. Both models publish open weights. On the current automated evidence, DeepSeek V4.1 Flash (max) is the stronger default; MiniMax M3 remains a defensible choice when its specific strengths, price, or deployment profile matter more than the headline rank.
Pick DeepSeek V4.1 Flash (max) when you prioritize the stronger current board result, slop score (eq-bench + ours), youtube best practices, open weights, and lower measured cost.
Pick MiniMax M3 when you prioritize open weights.
Blue bars: DeepSeek V4.1 Flash (max). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| DeepSeek V4.1 Flash (max) | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 86.8 | 83.4 |
| Writing Elo | 2085 | 1760 |
| Run-to-run spread (± overall std) | 6.940 | 2.030 |
| Cost per script (USD) | 0.012 | 0.023 |
| Avg latency (s) | 74.6 | 109.9 |
| Open weights | Yes | Yes |
Full scorecards: DeepSeek V4.1 Flash (max) · MiniMax M3. How scoring works: methodology.
← All comparisons