Towards AITowards AIToneBench

MiniMax M3 vs GLM-5

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 82.4. Their confidence intervals overlap, so treat the order as close rather than settled.

MiniMax M3
#42 Elo 1790 · 82.4/100
GLM-5
#43 Elo 1783 · 82.4/100
Cost / script
$0.019 vs $0.023
Human baseline
92.7 MiniMax M3 falls below it · GLM-5 falls below it

The verdict

Two open models, basically the same price at about 2 cents a script, and the honest answer is I can't call a decisive winner. They land in a dead heat, 82.4 overall apiece, and the Elo intervals overlap, so there is no score verdict to hand out. What I can tell you is where each one falls apart. GLM-5's weak spot is numbers: 75.7 on our anti-slop check, which usually means reciting stats instead of reframing them the way I would on camera. MiniMax M3's problem is consistency. Its overall spread is about half again as wide as GLM's, and its task scores swing from 75.5 on one topic to 87 on another. One script lands, the next loses the YouTube structure, and it averages just 77.7 there. Both are genuinely usable open options, and cost won't decide anything for you here. If I had to ship weekly without reviewing every draft, I'd take the model that fails predictably. That's GLM-5.

Pick GLM-5 if you want the more consistent of two cheap open models and can clean up how it handles numbers.
Pick MiniMax M3 if you review every draft anyway and your topics land on its strong side, because its quality swings hard from task to task.

Metric by metric

Blue bars: MiniMax M3. Orange bars: GLM-5. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.8
83.1
Writing Craft & Clarity13% weight
83.9
84.0
Substance, Accuracy & Value15% weight
81.8
84.0
Continuity & Emotion14% weight
79.2
79.7
YouTube Best Practices12% weight
77.7
78.3
Hook Strength10% weight
86.1
86.5
Length Adherence8% weight
82.6
78.6
Slop Score (EQ-Bench + ours)5% weight
87.3
85.3
Visual Cue Quality4% weight
83.0
83.8

Everything else that differs

MiniMax M3GLM-5
Overall / 10082.482.4
Writing Elo17901783
Run-to-run spread (± overall std)5.6103.760
Cost per script (USD)0.0190.023
Avg latency (s)185.4117.3
Open weightsYesYes

Full scorecards: MiniMax M3 · GLM-5. How scoring works: methodology.

← All comparisons