Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 82.2.
This one is not close. Grok 4.6 (high) sits 17th at 2144.5 Elo; MiniMax M3 sits at #51 at 1734.6, a gap of 410 Elo points. The confidence intervals, 2103 to 2181 against 1675 to 1787, are nowhere near each other, so the gap is real. On raw scores it is 88.14 to 82.19, and MiniMax is also less consistent, with a standard deviation of 5.22 versus 1.43. The widest single gap is on numbers and slop discipline: 88.24 to 78.24 on the anti-slop metric, a full ten points. YouTube best practices is next, 87.08 to 78.03. MiniMax's one bright spot is a marginally better slop score, 95.49 to 95.44, which changes nothing. The case for MiniMax is entirely practical: it is open weights and costs about two cents per script versus about twelve cents for Grok, both exact figures. Cheap drafts, but you will be editing them.
Pick Grok 4.6 (high) if you want consistent, disciplined scripts that need little editing and about twelve cents per script is acceptable.
Pick MiniMax M3 if you want open weights and drafts at about two cents per script, and you are prepared to edit for consistency and slop.
Blue bars: Grok 4.6 (high). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| Grok 4.6 (high) | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 88.1 | 82.2 |
| Writing Elo | 2144 | 1735 |
| Run-to-run spread (± overall std) | 1.430 | 5.220 |
| Cost per script (USD) | 0.115 | 0.020 |
| Avg latency (s) | 199.2 | 201.6 |
| Open weights | No | Yes |
Full scorecards: Grok 4.6 (high) · MiniMax M3. How scoring works: methodology.
← All comparisons