Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (high) leads overall, 88.4 to 82.4.
On averages, GPT-5.6 wins comfortably, around 88 overall to MiniMax's 82. But the number I keep coming back to is the spread: MiniMax's overall standard deviation is 5.61, nearly four times GPT's, and its length adherence swings by almost 13 points run to run. Translation: one MiniMax script reads fine, the next blows past the target length and loses the plot. Its worst task dipped to 75.5 while GPT-5.6 never dropped below about 87 on anything. The structural side, retention beats, CTAs, section flow, is where MiniMax bleeds most: 77.7 on YouTube practices against GPT's 86.5. To be fair, MiniMax is open, its hooks are actually decent, and at an estimated two cents a script it's about an eighth of GPT's cost. If you're batch-generating drafts and cherry-picking, that's a workable trade. But if a script goes anywhere near publishing without heavy edits, I wouldn't gamble on that variance. Neither model touches the human baseline of about 96, by the way. GPT-5.6, without much hesitation.
Pick GPT-5.6 Sol if you need a consistent, structurally sound draft every single run.
Pick MiniMax M3 if you're batch-generating cheap open-weights drafts and plan to keep the good runs and toss the rest.
Blue bars: GPT-5.6 Sol (high). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (high) | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 88.4 | 82.4 |
| Writing Elo | 2315 | 1790 |
| Run-to-run spread (± overall std) | 1.750 | 5.610 |
| Cost per script (USD) | 0.163 | 0.019 |
| Avg latency (s) | 106.4 | 185.4 |
| Open weights | No | Yes |
Full scorecards: GPT-5.6 Sol (high) · MiniMax M3. How scoring works: methodology.
← All comparisons