Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.8 Max leads overall, 83.4 to 82.4. Their confidence intervals overlap, so treat the order as close rather than settled.
Qwen3.8 Max at #38 and MiniMax M3 at #42, 83.12 against 82.43. Just over one and a half points. Qwen's advantage is concentrated in length adherence, 92.90 against 82.63, and anti-slop, 87.78 against 79.01. MiniMax takes voice 83.78 against 81.93 and hooks 86.12 against 83.50, so it reads more naturally while being looser about the brief. Both are weak on continuity, 78.08 and 79.16, effectively tied and both poor: neither carries a long script cleanly by itself. Practically MiniMax is under two cents a script against Qwen's 15, so eight times cheaper, it is open weights against closed, and it is more than twice as fast at 185 seconds against 397. For a point and a half, the cheaper open model is the easier call unless you specifically need Qwen's length control.
Pick Qwen3.8 Max when the word count and slop discipline have to be right.
Pick MiniMax M3 for open weights at an eighth the cost and twice the speed, with better voice and hooks.
Blue bars: Qwen3.8 Max. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| Qwen3.8 Max | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 83.4 | 82.4 |
| Writing Elo | 1866 | 1790 |
| Run-to-run spread (± overall std) | 5.620 | 5.610 |
| Cost per script (USD) | 0.150 | 0.019 |
| Avg latency (s) | 397.4 | 185.4 |
| Open weights | No | Yes |
Full scorecards: Qwen3.8 Max · MiniMax M3. How scoring works: methodology.
← All comparisons