Towards AITowards AIToneBench

Grok 4.5 vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 82.4.

Grok 4.5
#30 Elo 2042 · 85.8/100
MiniMax M3
#42 Elo 1790 · 82.4/100
Cost / script
$0.038 vs $0.019
Human baseline
92.7 Grok 4.5 falls below it · MiniMax M3 falls below it

The verdict

Same price, roughly two cents a script either way, so this comes down purely to the writing, and Grok 4.5 takes it. The overall gap is 85.8 versus 81.8, but the honest read is about reliability more than averages. MiniMax's overall standard deviation is 5.61, and its Elo interval stretches across about 155 points, so I'm genuinely unsure how good it is on any given run. Grok's band sits entirely above it, which settles the ranking even if it flatters nobody. The clearest concrete gap is structure: 83.8 versus 77.7 on YouTube best practices, so MiniMax drafts tend to need their retention beats rebuilt. Both are shaky on length, high 70s to mid 80s with big spreads, so neither escapes an edit pass there. Points for MiniMax: it's open weights, and its hooks hold up fine. Grok also answers in under 30 seconds versus over two and a half minutes, which adds up when you iterate. Predictable output at this price is Grok's case. Open weights at this price is MiniMax's, and honestly, it's a real one.

Pick Grok 4.5 if you want the fast, predictable, better-structured script at this price.
Pick MiniMax M3 if you need open weights this cheap and can live with the run-to-run swings.

Metric by metric

Blue bars: Grok 4.5. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.1
83.8
Writing Craft & Clarity13% weight
86.6
83.9
Substance, Accuracy & Value15% weight
87.4
81.8
Continuity & Emotion14% weight
83.5
79.2
YouTube Best Practices12% weight
84.4
77.7
Hook Strength10% weight
88.0
86.1
Length Adherence8% weight
78.7
82.6
Slop Score (EQ-Bench + ours)5% weight
90.2
87.3
Visual Cue Quality4% weight
86.5
83.0

Everything else that differs

Grok 4.5MiniMax M3
Overall / 10085.882.4
Writing Elo20421790
Run-to-run spread (± overall std)2.8105.610
Cost per script (USD)0.0380.019
Avg latency (s)41.7185.4
Open weightsNoYes

Full scorecards: Grok 4.5 · MiniMax M3. How scoring works: methodology.

← All comparisons