Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 83.4 to 81.1.
MiniMax M3 sits at #42 and Qwen3.8 Max at #46, 83.44 against 81.06, so a gap of 2.38 on the overall score. The Elo gap is 1781.0 to 1713.9, but the confidence intervals still overlap at the edges, so the ranking is suggestive rather than settled. MiniMax leads across most of the board: voice, 83.34 to 80.65, writing quality, 84.38 to 81.05, hooks, 86.51 to 83.39, and even length adherence, 89.72 against 86.34. It is also far steadier, with an overall spread of 2.03 against Qwen's 7.89. Qwen answers on anti-slop with numbers, 86.27 against 78.11, and on the slop metric, 90.43 to 86.04, so it stays cleaner around figures while losing on nearly everything else. Both are weak on continuity, 80.48 and 76.22, with Qwen clearly worse: neither carries a long script cleanly on its own. Then the practical part. MiniMax costs about two cents per script against Qwen's about sixteen cents, roughly 7x cheaper, and it is open weights against a closed model. When the cheaper open model also scores higher and more consistently, it is the easy call unless you specifically need Qwen's handling of numbers.
Pick Qwen3.8 Max if length adherence and clean handling of numbers matter more than price and you want the higher-ranked model even at roughly 7x the cost.
Pick MiniMax M3 if you want open weights and near-equal overall quality at about two cents per script, with stronger hooks and a more natural voice.
Blue bars: MiniMax M3. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| MiniMax M3 | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 83.4 | 81.1 |
| Writing Elo | 1781 | 1714 |
| Run-to-run spread (± overall std) | 2.030 | 7.890 |
| Cost per script (USD) | 0.023 | 0.164 |
| Avg latency (s) | 109.9 | 429.0 |
| Open weights | Yes | No |
Full scorecards: MiniMax M3 · Qwen3.8 Max. How scoring works: methodology.
← All comparisons