Towards AITowards AIToneBench

MiniMax M3 vs Qwen3.8 Max

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. MiniMax M3 leads overall, 83.4 to 81.1.

MiniMax M3
#42 Elo 1781 · 83.4/100
Qwen3.8 Max
#46 Elo 1714 · 81.1/100
Cost / script
$0.023 vs $0.164
Human baseline
90.2 MiniMax M3 falls below it · Qwen3.8 Max falls below it

The verdict

MiniMax M3 sits at #42 and Qwen3.8 Max at #46, 83.44 against 81.06, so a gap of 2.38 on the overall score. The Elo gap is 1781.0 to 1713.9, but the confidence intervals still overlap at the edges, so the ranking is suggestive rather than settled. MiniMax leads across most of the board: voice, 83.34 to 80.65, writing quality, 84.38 to 81.05, hooks, 86.51 to 83.39, and even length adherence, 89.72 against 86.34. It is also far steadier, with an overall spread of 2.03 against Qwen's 7.89. Qwen answers on anti-slop with numbers, 86.27 against 78.11, and on the slop metric, 90.43 to 86.04, so it stays cleaner around figures while losing on nearly everything else. Both are weak on continuity, 80.48 and 76.22, with Qwen clearly worse: neither carries a long script cleanly on its own. Then the practical part. MiniMax costs about two cents per script against Qwen's about sixteen cents, roughly 7x cheaper, and it is open weights against a closed model. When the cheaper open model also scores higher and more consistently, it is the easy call unless you specifically need Qwen's handling of numbers.

Pick Qwen3.8 Max if length adherence and clean handling of numbers matter more than price and you want the higher-ranked model even at roughly 7x the cost.
Pick MiniMax M3 if you want open weights and near-equal overall quality at about two cents per script, with stronger hooks and a more natural voice.

Metric by metric

Blue bars: MiniMax M3. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
83.3
80.7
Writing Craft & Clarity13% weight
84.4
81.0
Substance, Accuracy & Value15% weight
82.2
81.3
Continuity & Emotion14% weight
80.5
76.2
YouTube Best Practices12% weight
80.4
77.8
Hook Strength10% weight
86.5
83.4
Length Adherence8% weight
89.7
86.3
Slop Score (EQ-Bench + ours)5% weight
86.0
90.4
Visual Cue Quality4% weight
81.2
80.6

Everything else that differs

MiniMax M3Qwen3.8 Max
Overall / 10083.481.1
Writing Elo17811714
Run-to-run spread (± overall std)2.0307.890
Cost per script (USD)0.0230.164
Avg latency (s)109.9429.0
Open weightsYesNo

Full scorecards: MiniMax M3 · Qwen3.8 Max. How scoring works: methodology.

← All comparisons