Towards AITowards AIToneBench

Qwen3.8 Max vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.8 Max leads overall, 83.4 to 82.4. Their confidence intervals overlap, so treat the order as close rather than settled.

Qwen3.8 Max
#38 Elo 1866 · 83.4/100
MiniMax M3
#42 Elo 1790 · 82.4/100
Cost / script
$0.150 vs $0.019
Human baseline
92.7 Qwen3.8 Max falls below it · MiniMax M3 falls below it

The verdict

Qwen3.8 Max at #38 and MiniMax M3 at #42, 83.12 against 82.43. Just over one and a half points. Qwen's advantage is concentrated in length adherence, 92.90 against 82.63, and anti-slop, 87.78 against 79.01. MiniMax takes voice 83.78 against 81.93 and hooks 86.12 against 83.50, so it reads more naturally while being looser about the brief. Both are weak on continuity, 78.08 and 79.16, effectively tied and both poor: neither carries a long script cleanly by itself. Practically MiniMax is under two cents a script against Qwen's 15, so eight times cheaper, it is open weights against closed, and it is more than twice as fast at 185 seconds against 397. For a point and a half, the cheaper open model is the easier call unless you specifically need Qwen's length control.

Pick Qwen3.8 Max when the word count and slop discipline have to be right.
Pick MiniMax M3 for open weights at an eighth the cost and twice the speed, with better voice and hooks.

Metric by metric

Blue bars: Qwen3.8 Max. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
81.9
83.8
Writing Craft & Clarity13% weight
83.4
83.9
Substance, Accuracy & Value15% weight
84.4
81.8
Continuity & Emotion14% weight
78.1
79.2
YouTube Best Practices12% weight
80.4
77.7
Hook Strength10% weight
83.5
86.1
Length Adherence8% weight
92.9
82.6
Slop Score (EQ-Bench + ours)5% weight
92.0
87.3
Visual Cue Quality4% weight
85.2
83.0

Everything else that differs

Qwen3.8 MaxMiniMax M3
Overall / 10083.482.4
Writing Elo18661790
Run-to-run spread (± overall std)5.6205.610
Cost per script (USD)0.1500.019
Avg latency (s)397.4185.4
Open weightsNoYes

Full scorecards: Qwen3.8 Max · MiniMax M3. How scoring works: methodology.

← All comparisons