Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 82.4.
Same price, roughly two cents a script either way, so this comes down purely to the writing, and Grok 4.5 takes it. The overall gap is 85.8 versus 81.8, but the honest read is about reliability more than averages. MiniMax's overall standard deviation is 5.61, and its Elo interval stretches across about 155 points, so I'm genuinely unsure how good it is on any given run. Grok's band sits entirely above it, which settles the ranking even if it flatters nobody. The clearest concrete gap is structure: 83.8 versus 77.7 on YouTube best practices, so MiniMax drafts tend to need their retention beats rebuilt. Both are shaky on length, high 70s to mid 80s with big spreads, so neither escapes an edit pass there. Points for MiniMax: it's open weights, and its hooks hold up fine. Grok also answers in under 30 seconds versus over two and a half minutes, which adds up when you iterate. Predictable output at this price is Grok's case. Open weights at this price is MiniMax's, and honestly, it's a real one.
Pick Grok 4.5 if you want the fast, predictable, better-structured script at this price.
Pick MiniMax M3 if you need open weights this cheap and can live with the run-to-run swings.
Blue bars: Grok 4.5. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| Grok 4.5 | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 85.8 | 82.4 |
| Writing Elo | 2042 | 1790 |
| Run-to-run spread (± overall std) | 2.810 | 5.610 |
| Cost per script (USD) | 0.038 | 0.019 |
| Avg latency (s) | 41.7 | 185.4 |
| Open weights | No | Yes |
Full scorecards: Grok 4.5 · MiniMax M3. How scoring works: methodology.
← All comparisons