Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 81.7. Their confidence intervals overlap, so treat the order as close rather than settled.
MiniMax M3 at #42 and Qwen3.7 Max at #50, 82.43 against 81.73. Two thirds of a point, which on this board is a coin flip. MiniMax has the better voice and hooks, 83.80 and 86.12 against 81.15 and 84.25. Qwen answers on the structural metrics: length adherence 85.79 against 82.6, and YouTube best practices 80.82 against 77.73. So MiniMax writes the better sentences and Qwen builds the better shape. The practical differences are bigger than the quality ones. MiniMax is open weights at under two cents a script. Qwen is closed at five and a half. Nearly three times the price for what is, on this evidence, the same quality. Unless you need Qwen specifically, the open cheaper model is the easier call.
Pick MiniMax M3 if you want open weights at the lowest price here, with better voice and hooks.
Pick Qwen3.7 Max if you need better length control and structural discipline, or you're already on Alibaba infrastructure.
Blue bars: MiniMax M3. Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.
| MiniMax M3 | Qwen3.7 Max (default) | |
|---|---|---|
| Overall / 100 | 82.4 | 81.7 |
| Writing Elo | 1790 | 1699 |
| Run-to-run spread (± overall std) | 5.610 | 3.060 |
| Cost per script (USD) | 0.019 | 0.054 |
| Avg latency (s) | 185.4 | 156.8 |
| Open weights | Yes | No |
Full scorecards: MiniMax M3 · Qwen3.7 Max (default). How scoring works: methodology.
← All comparisons