Towards AITowards AIToneBench

Claude Opus 5 (max) vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 82.4.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
MiniMax M3
#42 Elo 1790 · 82.4/100
Cost / script
$0.120 vs $0.019
Human baseline
92.7 Claude Opus 5 (max) falls below it · MiniMax M3 falls below it

The verdict

MiniMax M3 is the cheapest model in this whole comparison, under two cents a script, open weights, #42 on the board. Against Opus 5 at max effort that's 82.43 overall against 90.80, so nearly nine points. The gap is not evenly spread. Its hooks are fine at 86.12 and its voice holds at 83.80, which is better than the ranking suggests. What breaks is structure: YouTube best practices at 77.73 is the worst metric in this pairing, over thirteen points below Opus, and continuity at 79.16 is twelve points down. So it opens well and then loses the thread. For a script, that's an expensive kind of failure, because the fix is structural rather than a line edit. At six times cheaper than Opus and fully open, it earns a place in a drafting pipeline where a human owns the outline and the model fills sections. It does not earn one where the model owns the arc.

Pick Opus 5 (max) if the model has to hold the whole script together, which is most real script work.
Pick MiniMax M3 if you own the structure yourself and want open weights filling in sections at under two cents a draft.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
83.8
Writing Craft & Clarity13% weight
91.1
83.9
Substance, Accuracy & Value15% weight
90.6
81.8
Continuity & Emotion14% weight
90.4
79.2
YouTube Best Practices12% weight
89.6
77.7
Hook Strength10% weight
92.0
86.1
Length Adherence8% weight
89.7
82.6
Slop Score (EQ-Bench + ours)5% weight
93.3
87.3
Visual Cue Quality4% weight
89.6
83.0

Everything else that differs

Claude Opus 5 (max)MiniMax M3
Overall / 10090.882.4
Writing Elo26491790
Run-to-run spread (± overall std)1.2905.610
Cost per script (USD)0.1200.019
Avg latency (s)243.8185.4
Open weightsNoYes

Full scorecards: Claude Opus 5 (max) · MiniMax M3. How scoring works: methodology.

← All comparisons