Towards AITowards AIToneBench

GPT-5.6 Sol (high) vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 82.4.

GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
MiniMax M3
#42 Elo 1790 · 82.4/100
Cost / script
$0.163 vs $0.019
Human baseline
92.7 GPT-5.6 Sol (high) falls below it · MiniMax M3 falls below it

The verdict

On averages, GPT-5.6 wins comfortably, around 88 overall to MiniMax's 82. But the number I keep coming back to is the spread: MiniMax's overall standard deviation is 5.61, nearly four times GPT's, and its length adherence swings by almost 13 points run to run. Translation: one MiniMax script reads fine, the next blows past the target length and loses the plot. Its worst task dipped to 75.5 while GPT-5.6 never dropped below about 87 on anything. The structural side, retention beats, CTAs, section flow, is where MiniMax bleeds most: 77.7 on YouTube practices against GPT's 86.5. To be fair, MiniMax is open, its hooks are actually decent, and at an estimated two cents a script it's about an eighth of GPT's cost. If you're batch-generating drafts and cherry-picking, that's a workable trade. But if a script goes anywhere near publishing without heavy edits, I wouldn't gamble on that variance. Neither model touches the human baseline of about 96, by the way. GPT-5.6, without much hesitation.

Pick GPT-5.6 Sol if you need a consistent, structurally sound draft every single run.
Pick MiniMax M3 if you're batch-generating cheap open-weights drafts and plan to keep the good runs and toss the rest.

Metric by metric

Blue bars: GPT-5.6 Sol (high). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.3
83.8
Writing Craft & Clarity13% weight
88.1
83.9
Substance, Accuracy & Value15% weight
90.1
81.8
Continuity & Emotion14% weight
86.1
79.2
YouTube Best Practices12% weight
86.5
77.7
Hook Strength10% weight
87.5
86.1
Length Adherence8% weight
91.7
82.6
Slop Score (EQ-Bench + ours)5% weight
93.4
87.3
Visual Cue Quality4% weight
91.0
83.0

Everything else that differs

GPT-5.6 Sol (high)MiniMax M3
Overall / 10088.482.4
Writing Elo23151790
Run-to-run spread (± overall std)1.7505.610
Cost per script (USD)0.1630.019
Avg latency (s)106.4185.4
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (high) · MiniMax M3. How scoring works: methodology.

← All comparisons