Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 81.5. Their confidence intervals overlap, so treat the order as close rather than settled.
This one is about as close to a coin flip as a head-to-head gets. MiniMax M3 edges the Elo (1790 vs 1674) and now the overall score too (82.43 vs 81.51), and the confidence intervals swallow both gaps completely. So forget the ranks; the real difference is temperament. MiniMax is the higher-ceiling, higher-variance writer: a better tone match at 83.8 vs 80.7, but an overall spread roughly three times wider than Qwen's, with task scores swinging from 75.5 to 87 depending on the topic. Qwen is the steady one: 83.3 on length adherence against MiniMax's 82.6, and it holds YouTube structure noticeably better. Neither nails my voice, to be clear; both sit about 12 points under the human baseline. Cost leans MiniMax, about 2 cents a script versus 5, both estimates, and it's open weights if you want to self-host. My honest read: choose by workflow, not by rank. If you review every draft anyway, take MiniMax's ceiling. If you're automating, take Qwen's floor.
Pick MiniMax M3 if you're reviewing drafts by hand and want the better voice match at the lower price, open weights included.
Pick Qwen3.7 Max if you're automating and need predictable length and structure more than a higher ceiling.
Blue bars: MiniMax M3. Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.
| MiniMax M3 | Qwen3.7 Max (high) | |
|---|---|---|
| Overall / 100 | 82.4 | 81.5 |
| Writing Elo | 1790 | 1674 |
| Run-to-run spread (± overall std) | 5.610 | 2.580 |
| Cost per script (USD) | 0.019 | 0.054 |
| Avg latency (s) | 185.4 | 157.2 |
| Open weights | Yes | No |
Full scorecards: MiniMax M3 · Qwen3.7 Max (high). How scoring works: methodology.
← All comparisons