Towards AITowards AIToneBench

Grok 4.6 (high) vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 82.2.

Grok 4.6 (high)
#18 Elo 2144 · 88.1/100
MiniMax M3
#46 Elo 1735 · 82.2/100
Cost / script
$0.115 vs $0.020
Human baseline
90.3 Grok 4.6 (high) falls below it · MiniMax M3 falls below it

The verdict

This one is not close. Grok 4.6 (high) sits 17th at 2144.5 Elo; MiniMax M3 sits at #51 at 1734.6, a gap of 410 Elo points. The confidence intervals, 2103 to 2181 against 1675 to 1787, are nowhere near each other, so the gap is real. On raw scores it is 88.14 to 82.19, and MiniMax is also less consistent, with a standard deviation of 5.22 versus 1.43. The widest single gap is on numbers and slop discipline: 88.24 to 78.24 on the anti-slop metric, a full ten points. YouTube best practices is next, 87.08 to 78.03. MiniMax's one bright spot is a marginally better slop score, 95.49 to 95.44, which changes nothing. The case for MiniMax is entirely practical: it is open weights and costs about two cents per script versus about twelve cents for Grok, both exact figures. Cheap drafts, but you will be editing them.

Pick Grok 4.6 (high) if you want consistent, disciplined scripts that need little editing and about twelve cents per script is acceptable.
Pick MiniMax M3 if you want open weights and drafts at about two cents per script, and you are prepared to edit for consistency and slop.

Metric by metric

Blue bars: Grok 4.6 (high). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
88.8
83.4
Writing Craft & Clarity13% weight
88.0
83.5
Substance, Accuracy & Value15% weight
88.0
81.6
Continuity & Emotion14% weight
86.3
78.5
YouTube Best Practices12% weight
87.1
78.0
Hook Strength10% weight
90.0
86.0
Length Adherence8% weight
88.3
83.1
Slop Score (EQ-Bench + ours)5% weight
91.8
86.9
Visual Cue Quality4% weight
86.4
82.6

Everything else that differs

Grok 4.6 (high)MiniMax M3
Overall / 10088.182.2
Writing Elo21441735
Run-to-run spread (± overall std)1.4305.220
Cost per script (USD)0.1150.020
Avg latency (s)199.2201.6
Open weightsNoYes

Full scorecards: Grok 4.6 (high) · MiniMax M3. How scoring works: methodology.

← All comparisons