Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 82.4.
MiniMax M3 is the cheapest model in this whole comparison, under two cents a script, open weights, #42 on the board. Against Opus 5 at max effort that's 82.43 overall against 90.80, so nearly nine points. The gap is not evenly spread. Its hooks are fine at 86.12 and its voice holds at 83.80, which is better than the ranking suggests. What breaks is structure: YouTube best practices at 77.73 is the worst metric in this pairing, over thirteen points below Opus, and continuity at 79.16 is twelve points down. So it opens well and then loses the thread. For a script, that's an expensive kind of failure, because the fix is structural rather than a line edit. At six times cheaper than Opus and fully open, it earns a place in a drafting pipeline where a human owns the outline and the model fills sections. It does not earn one where the model owns the arc.
Pick Opus 5 (max) if the model has to hold the whole script together, which is most real script work.
Pick MiniMax M3 if you own the structure yourself and want open weights filling in sections at under two cents a draft.
Blue bars: Claude Opus 5 (max). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 90.8 | 82.4 |
| Writing Elo | 2649 | 1790 |
| Run-to-run spread (± overall std) | 1.290 | 5.610 |
| Cost per script (USD) | 0.120 | 0.019 |
| Avg latency (s) | 243.8 | 185.4 |
| Open weights | No | Yes |
Full scorecards: Claude Opus 5 (max) · MiniMax M3. How scoring works: methodology.
← All comparisons