Towards AITowards AIToneBench

GPT-5.6 Sol (ultra) vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.4 to 82.4.

GPT-5.6 Sol (ultra)
#16 Elo 2316 · 88.4/100
MiniMax M3
#42 Elo 1790 · 82.4/100
Cost / script
$0.137 vs $0.019
Human baseline
92.7 GPT-5.6 Sol (ultra) falls below it · MiniMax M3 falls below it

The verdict

GPT-5.6 Sol at ultra is #16 against MiniMax M3 at #42, 88.40 overall to 82.43. Six and a half points. The widest gap is YouTube structure, 86.17 against 77.73, and length adherence is close behind, 94.34 against 82.63. Both describe the same thing: MiniMax loses the shape of a long script. Its hooks are fine at 86.12 and its voice at 83.80 is better than the ranking suggests, so the sentences are not the problem, the architecture is. The economics are stark: under two cents a script against 14 cents, so seven times cheaper, open weights, and nearly three times faster at 185 seconds against 421. That makes MiniMax a reasonable section-filler when a human owns the outline, and a poor choice when the model owns the arc.

Pick GPT-5.6 Sol (ultra) when the model has to hold the whole structure, which is most real script work.
Pick MiniMax M3 if you own the outline and want open weights filling sections, seven times cheaper and much faster.

Metric by metric

Blue bars: GPT-5.6 Sol (ultra). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
87.0
83.8
Writing Craft & Clarity13% weight
88.1
83.9
Substance, Accuracy & Value15% weight
90.5
81.8
Continuity & Emotion14% weight
85.7
79.2
YouTube Best Practices12% weight
86.1
77.7
Hook Strength10% weight
86.4
86.1
Length Adherence8% weight
94.3
82.6
Slop Score (EQ-Bench + ours)5% weight
93.5
87.3
Visual Cue Quality4% weight
91.2
83.0

Everything else that differs

GPT-5.6 Sol (ultra)MiniMax M3
Overall / 10088.482.4
Writing Elo23161790
Run-to-run spread (± overall std)1.9705.610
Cost per script (USD)0.1370.019
Avg latency (s)420.9185.4
Open weightsNoYes

Full scorecards: GPT-5.6 Sol (ultra) · MiniMax M3. How scoring works: methodology.

← All comparisons