Towards AITowards AIToneBench

Grok 4.6 vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 leads overall, 86.6 to 83.4.

Grok 4.6
#26 Elo 2035 · 86.6/100
MiniMax M3
#42 Elo 1781 · 83.4/100
Cost / script
$0.204 vs $0.023
Human baseline
90.2 Grok 4.6 falls below it · MiniMax M3 falls below it

The verdict

Grok 4.6 wins this one comfortably. It sits at #26 with 2035.3 Elo against MiniMax M3 at #42 with 1781.0, and the confidence intervals don't come close to touching, so the gap is real. Overall it's 86.61 to 83.44, a 3.17-point spread, and Grok takes most of the metrics I care about: tone and voice 87.86 vs 83.34, substance 87.53 vs 82.23, continuity and emotion 85.78 vs 80.48, and the numbers-and-slop metric by a wide 87.01 to 78.11. MiniMax pushes back in two places: hooks, where it clearly wins 86.51 to Grok's 80.81, and length adherence, 89.72 to 88.03. The records agree with the ranking, 941 wins and 99 losses for Grok against MiniMax's 737 and 294. Then the bill flips the conversation. MiniMax is open weights and costs about two cents per script against roughly twenty cents for Grok, close to 9x cheaper. If the script ships close to as-written, pay for Grok. If it's a volume pass you'll rewrite anyway, the two-cent draft is hard to ignore.

Pick Grok 4.6 if you want the clearly stronger writer on tone, substance, and continuity and can pay roughly twenty cents per script for it.
Pick MiniMax M3 if you want open weights at about two cents a script with stronger hooks, and can accept a 3.17-point overall gap on drafts you plan to edit anyway.

Metric by metric

Blue bars: Grok 4.6. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.9
83.3
Writing Craft & Clarity13% weight
86.5
84.4
Substance, Accuracy & Value15% weight
87.5
82.2
Continuity & Emotion14% weight
85.8
80.5
YouTube Best Practices12% weight
86.5
80.4
Hook Strength10% weight
80.8
86.5
Length Adherence8% weight
88.0
89.7
Slop Score (EQ-Bench + ours)5% weight
91.2
86.0
Visual Cue Quality4% weight
86.6
81.2

Everything else that differs

Grok 4.6MiniMax M3
Overall / 10086.683.4
Writing Elo20351781
Run-to-run spread (± overall std)2.5302.030
Cost per script (USD)0.2040.023
Avg latency (s)243.1109.9
Open weightsNoYes

Full scorecards: Grok 4.6 · MiniMax M3. How scoring works: methodology.

← All comparisons