Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. GPT-5.6 Sol (ultra) leads overall, 88.4 to 82.4.
GPT-5.6 Sol at ultra is #16 against MiniMax M3 at #42, 88.40 overall to 82.43. Six and a half points. The widest gap is YouTube structure, 86.17 against 77.73, and length adherence is close behind, 94.34 against 82.63. Both describe the same thing: MiniMax loses the shape of a long script. Its hooks are fine at 86.12 and its voice at 83.80 is better than the ranking suggests, so the sentences are not the problem, the architecture is. The economics are stark: under two cents a script against 14 cents, so seven times cheaper, open weights, and nearly three times faster at 185 seconds against 421. That makes MiniMax a reasonable section-filler when a human owns the outline, and a poor choice when the model owns the arc.
Pick GPT-5.6 Sol (ultra) when the model has to hold the whole structure, which is most real script work.
Pick MiniMax M3 if you own the outline and want open weights filling sections, seven times cheaper and much faster.
Blue bars: GPT-5.6 Sol (ultra). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| GPT-5.6 Sol (ultra) | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 88.4 | 82.4 |
| Writing Elo | 2316 | 1790 |
| Run-to-run spread (± overall std) | 1.970 | 5.610 |
| Cost per script (USD) | 0.137 | 0.019 |
| Avg latency (s) | 420.9 | 185.4 |
| Open weights | No | Yes |
Full scorecards: GPT-5.6 Sol (ultra) · MiniMax M3. How scoring works: methodology.
← All comparisons