Towards AITowards AIToneBench

MiniMax M3 vs Qwen3.7 Max (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 82.4 to 81.7. Their confidence intervals overlap, so treat the order as close rather than settled.

MiniMax M3
#42 Elo 1790 · 82.4/100
Qwen3.7 Max (default)
#50 Elo 1699 · 81.7/100
Cost / script
$0.019 vs $0.054
Human baseline
92.7 MiniMax M3 falls below it · Qwen3.7 Max (default) falls below it

The verdict

MiniMax M3 at #42 and Qwen3.7 Max at #50, 82.43 against 81.73. Two thirds of a point, which on this board is a coin flip. MiniMax has the better voice and hooks, 83.80 and 86.12 against 81.15 and 84.25. Qwen answers on the structural metrics: length adherence 85.79 against 82.6, and YouTube best practices 80.82 against 77.73. So MiniMax writes the better sentences and Qwen builds the better shape. The practical differences are bigger than the quality ones. MiniMax is open weights at under two cents a script. Qwen is closed at five and a half. Nearly three times the price for what is, on this evidence, the same quality. Unless you need Qwen specifically, the open cheaper model is the easier call.

Pick MiniMax M3 if you want open weights at the lowest price here, with better voice and hooks.
Pick Qwen3.7 Max if you need better length control and structural discipline, or you're already on Alibaba infrastructure.

Metric by metric

Blue bars: MiniMax M3. Orange bars: Qwen3.7 Max (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.8
81.2
Writing Craft & Clarity13% weight
83.9
81.7
Substance, Accuracy & Value15% weight
81.8
80.9
Continuity & Emotion14% weight
79.2
79.0
YouTube Best Practices12% weight
77.7
80.8
Hook Strength10% weight
86.1
84.2
Length Adherence8% weight
82.6
85.8
Slop Score (EQ-Bench + ours)5% weight
87.3
86.0
Visual Cue Quality4% weight
83.0
80.2

Everything else that differs

MiniMax M3Qwen3.7 Max (default)
Overall / 10082.481.7
Writing Elo17901699
Run-to-run spread (± overall std)5.6103.060
Cost per script (USD)0.0190.054
Avg latency (s)185.4156.8
Open weightsYesNo

Full scorecards: MiniMax M3 · Qwen3.7 Max (default). How scoring works: methodology.

← All comparisons