Towards AITowards AIToneBench

GLM-5 vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 82.4. Their confidence intervals overlap, so treat the order as close rather than settled.

GLM-5
#43 Elo 1783 · 82.4/100
MiniMax M3
#42 Elo 1790 · 82.4/100
Cost / script
$0.023 vs $0.019
Human baseline
92.7 GLM-5 falls below it · MiniMax M3 falls below it

The verdict

Two open models, basically the same price at about 2 cents a script, and the honest answer is I can't call a decisive winner. They land in a dead heat, 82.4 overall apiece, and the Elo intervals overlap, so there is no score verdict to hand out. What I can tell you is where each one falls apart. GLM-5's weak spot is numbers: 75.7 on our anti-slop check, which usually means reciting stats instead of reframing them the way I would on camera. MiniMax M3's problem is consistency. Its overall spread is about half again as wide as GLM's, and its task scores swing from 75.5 on one topic to 87 on another. One script lands, the next loses the YouTube structure, and it averages just 77.7 there. Both are genuinely usable open options, and cost won't decide anything for you here. If I had to ship weekly without reviewing every draft, I'd take the model that fails predictably. That's GLM-5.

Pick GLM-5 if you want the more consistent of two cheap open models and can clean up how it handles numbers.
Pick MiniMax M3 if you review every draft anyway and your topics land on its strong side, because its quality swings hard from task to task.

Metric by metric

Blue bars: GLM-5. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.1
83.8
Writing Craft & Clarity13% weight
84.0
83.9
Substance, Accuracy & Value15% weight
84.0
81.8
Continuity & Emotion14% weight
79.7
79.2
YouTube Best Practices12% weight
78.3
77.7
Hook Strength10% weight
86.5
86.1
Length Adherence8% weight
78.6
82.6
Slop Score (EQ-Bench + ours)5% weight
85.3
87.3
Visual Cue Quality4% weight
83.8
83.0

Everything else that differs

GLM-5MiniMax M3
Overall / 10082.482.4
Writing Elo17831790
Run-to-run spread (± overall std)3.7605.610
Cost per script (USD)0.0230.019
Avg latency (s)117.3185.4
Open weightsYesYes

Full scorecards: GLM-5 · MiniMax M3. How scoring works: methodology.

← All comparisons