Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 82.4. Their confidence intervals overlap, so treat the order as close rather than settled.
Two open models, basically the same price at about 2 cents a script, and the honest answer is I can't call a decisive winner. They land in a dead heat, 82.4 overall apiece, and the Elo intervals overlap, so there is no score verdict to hand out. What I can tell you is where each one falls apart. GLM-5's weak spot is numbers: 75.7 on our anti-slop check, which usually means reciting stats instead of reframing them the way I would on camera. MiniMax M3's problem is consistency. Its overall spread is about half again as wide as GLM's, and its task scores swing from 75.5 on one topic to 87 on another. One script lands, the next loses the YouTube structure, and it averages just 77.7 there. Both are genuinely usable open options, and cost won't decide anything for you here. If I had to ship weekly without reviewing every draft, I'd take the model that fails predictably. That's GLM-5.
Pick GLM-5 if you want the more consistent of two cheap open models and can clean up how it handles numbers.
Pick MiniMax M3 if you review every draft anyway and your topics land on its strong side, because its quality swings hard from task to task.
Blue bars: GLM-5. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| GLM-5 | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 82.4 | 82.4 |
| Writing Elo | 1783 | 1790 |
| Run-to-run spread (± overall std) | 3.760 | 5.610 |
| Cost per script (USD) | 0.023 | 0.019 |
| Avg latency (s) | 117.3 | 185.4 |
| Open weights | Yes | Yes |
Full scorecards: GLM-5 · MiniMax M3. How scoring works: methodology.
← All comparisons