Towards AITowards AIToneBench

GLM-5.3 vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GLM-5.3 leads overall, 86.8 to 83.4.

GLM-5.3
#24 Elo 2042 · 86.8/100
MiniMax M3
#42 Elo 1781 · 83.4/100
Cost / script
$0.109 vs $0.023
Human baseline
90.2 GLM-5.3 falls below it · MiniMax M3 falls below it

The verdict

GLM-5.3 wins this one, and the intervals say it's real: its Elo floor of 1989.0 sits well above MiniMax M3's ceiling of 1813.1. On the board that's #24 at 2041.5 Elo against #42 at 1781.0, a 260.5-point gap, and 86.83 to 83.44 overall. The records make it concrete: GLM took 866 wins with just 8 losses, while MiniMax lost 294 times. The quality gap is widest exactly where scripts go wrong. The numbers-and-slop metric is 86.75 to 78.11, slop is 91.64 to 86.04, and GLM also takes tone and voice, 87.89 vs 83.34, and continuity and emotion, 85.11 vs 80.48. MiniMax gets two things back: length adherence, 89.72 to GLM's 87.5, and an essential tie on visual cues, 81.21 to 81.43. Then the bill. Both are open weights, but MiniMax costs about two cents per script against GLM's eleven, nearly 5x cheaper, and it turns scripts around far faster, 109.94 seconds on average against GLM's 301.95. That trade is live if the drafts get rewritten anyway; if they're supposed to ship close to final, the 3.39-point overall gap is where you'll feel it.

Pick GLM-5.3 if you want the clearly stronger writer, with better tone, cleaner numbers, and far less slop, at about eleven cents per script.
Pick MiniMax M3 if you want scripts at about two cents each with much faster turnaround and can live with a 3.39-point overall gap on drafts you'll edit anyway.

Metric by metric

Blue bars: GLM-5.3. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.9
83.3
Writing Craft & Clarity13% weight
87.2
84.4
Substance, Accuracy & Value15% weight
85.9
82.2
Continuity & Emotion14% weight
85.1
80.5
YouTube Best Practices12% weight
85.0
80.4
Hook Strength10% weight
89.5
86.5
Length Adherence8% weight
87.5
89.7
Slop Score (EQ-Bench + ours)5% weight
91.6
86.0
Visual Cue Quality4% weight
81.4
81.2

Everything else that differs

GLM-5.3MiniMax M3
Overall / 10086.883.4
Writing Elo20421781
Run-to-run spread (± overall std)6.6002.030
Cost per script (USD)0.1090.023
Avg latency (s)301.9109.9
Open weightsYesYes

Full scorecards: GLM-5.3 · MiniMax M3. How scoring works: methodology.

← All comparisons