Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. DeepSeek V4 Flash 0731 leads overall, 84.2 to 82.4. Their confidence intervals overlap, so treat the order as close rather than settled.
DeepSeek V4 Flash 0731 at #34 and MiniMax M3 at #42, 84.23 against 82.43. Just under two points, and DeepSeek wins nearly every column: hooks 87.55 to 86.12, voice 85.44 to 83.80, continuity 82.02 to 79.16, YouTube structure 81.06 to 77.73, number discipline 81.75 to 79.01, even length adherence 83.57 to 82.63. MiniMax's remaining edges are cue quality and speed. Cost: half a cent against just under two cents, so DeepSeek is roughly four times cheaper as well. MiniMax is noticeably faster, 185 seconds against 289. Given the new build's 49-place jump, DeepSeek is the one I would default to here.
Pick DeepSeek V4 Flash 0731 for better hooks, voice and structure at a quarter the price.
Pick MiniMax M3 if you want the tighter length control of the two, or you need drafts back twice as fast.
Blue bars: DeepSeek V4 Flash 0731. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| DeepSeek V4 Flash 0731 | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 84.2 | 82.4 |
| Writing Elo | 1955 | 1790 |
| Run-to-run spread (± overall std) | 9.240 | 5.610 |
| Cost per script (USD) | 0.005 | 0.019 |
| Avg latency (s) | 289.0 | 185.4 |
| Open weights | Yes | Yes |
Full scorecards: DeepSeek V4 Flash 0731 · MiniMax M3. How scoring works: methodology.
← All comparisons