Towards AITowards AIToneBench

MiniMax M3 vs Qwen3.7 Max (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 81.5. Their confidence intervals overlap, so treat the order as close rather than settled.

MiniMax M3
#42 Elo 1790 · 82.4/100
Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
Cost / script
$0.019 vs $0.054
Human baseline
92.7 MiniMax M3 falls below it · Qwen3.7 Max (high) falls below it

The verdict

This one is about as close to a coin flip as a head-to-head gets. MiniMax M3 edges the Elo (1790 vs 1674) and now the overall score too (82.43 vs 81.51), and the confidence intervals swallow both gaps completely. So forget the ranks; the real difference is temperament. MiniMax is the higher-ceiling, higher-variance writer: a better tone match at 83.8 vs 80.7, but an overall spread roughly three times wider than Qwen's, with task scores swinging from 75.5 to 87 depending on the topic. Qwen is the steady one: 83.3 on length adherence against MiniMax's 82.6, and it holds YouTube structure noticeably better. Neither nails my voice, to be clear; both sit about 12 points under the human baseline. Cost leans MiniMax, about 2 cents a script versus 5, both estimates, and it's open weights if you want to self-host. My honest read: choose by workflow, not by rank. If you review every draft anyway, take MiniMax's ceiling. If you're automating, take Qwen's floor.

Pick MiniMax M3 if you're reviewing drafts by hand and want the better voice match at the lower price, open weights included.
Pick Qwen3.7 Max if you're automating and need predictable length and structure more than a higher ceiling.

Metric by metric

Blue bars: MiniMax M3. Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.8
80.9
Writing Craft & Clarity13% weight
83.9
81.5
Substance, Accuracy & Value15% weight
81.8
81.2
Continuity & Emotion14% weight
79.2
79.2
YouTube Best Practices12% weight
77.7
80.7
Hook Strength10% weight
86.1
85.0
Length Adherence8% weight
82.6
83.3
Slop Score (EQ-Bench + ours)5% weight
87.3
85.9
Visual Cue Quality4% weight
83.0
78.4

Everything else that differs

MiniMax M3Qwen3.7 Max (high)
Overall / 10082.481.5
Writing Elo17901674
Run-to-run spread (± overall std)5.6102.580
Cost per script (USD)0.0190.054
Avg latency (s)185.4157.2
Open weightsYesNo

Full scorecards: MiniMax M3 · Qwen3.7 Max (high). How scoring works: methodology.

← All comparisons