Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 80.2.
On the averages, MiniMax M3 leads: 82.6 vs 80.2 overall and roughly 210 Elo points ahead. But the intervals overlap, and MiniMax's overall spread is nearly four times wider than Gemini's, so I wouldn't bet much on that gap holding. Here's how I read it. When MiniMax lands, it sounds more like me: 83.8 vs 80.0 on tone and 84.0 vs 80.8 on writing craft. The catch is that it doesn't always land; it scored 75.5 on one of our tasks and about 87 on another. Gemini 3.1 Pro is flatter on voice but steadier, and it keeps YouTube structure better, 78.9 against MiniMax's 77.7, which is MiniMax's weakest metric here. Neither gets near the human baseline of 92.7, let's be honest. The practical split: MiniMax is open weights at about 2 cents a script, estimated, while Gemini costs six times that and is more than twice as fast per draft. Ceiling versus consistency, then. This time, cost and open weights tip me to MiniMax.
Pick MiniMax M3 if cost and open weights matter to you and you can absorb its script-to-script swings.
Pick Gemini 3.1 Pro if you need steadier output and faster drafts and can live with a flatter voice.
Blue bars: MiniMax M3. Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.
| MiniMax M3 | Gemini 3.1 Pro (default) | |
|---|---|---|
| Overall / 100 | 82.4 | 80.2 |
| Writing Elo | 1790 | 1577 |
| Run-to-run spread (± overall std) | 5.610 | 2.420 |
| Cost per script (USD) | 0.019 | 0.118 |
| Avg latency (s) | 185.4 | 63.8 |
| Open weights | Yes | No |
Full scorecards: MiniMax M3 · Gemini 3.1 Pro (default). How scoring works: methodology.
← All comparisons