Towards AITowards AIToneBench

Kimi K3 (thinking) vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 82.4.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
MiniMax M3
#42 Elo 1790 · 82.4/100
Cost / script
$0.187 vs $0.019
Human baseline
92.7 Kimi K3 (thinking) falls below it · MiniMax M3 falls below it

The verdict

The headline says Kimi K3 at 89.2 against MiniMax M3 at 81.8, but the number that actually decides this matchup is the variance. MiniMax's overall spread is 5.61 points, nearly four times Kimi's, and its length adherence swings hardest of any metric here. In plain English: sometimes it lands the target, sometimes it's way off, and you won't know until you read it. On my learn-AI-engineering task it dropped to 75.5 while Kimi stayed above 88 on every task. That unreliability matters more to me than the average, because a model I have to babysit isn't saving me time. Credit to MiniMax though: it's open weights, the hooks are respectable at 86.1, and at roughly 2 cents per script, estimated, it's a fraction of Kimi's cost. The Elo intervals don't overlap, so the ranking itself isn't in doubt. If you want cheap open-weights drafts and you'll review every one, MiniMax is fine. If you want to trust what comes out, Kimi.

Pick Kimi K3 if you need output you can trust run after run, across every task type.
Pick MiniMax M3 if you want very cheap open-weights drafts and plan to review every single one.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
83.8
Writing Craft & Clarity13% weight
89.3
83.9
Substance, Accuracy & Value15% weight
89.2
81.8
Continuity & Emotion14% weight
88.3
79.2
YouTube Best Practices12% weight
87.8
77.7
Hook Strength10% weight
90.2
86.1
Length Adherence8% weight
91.1
82.6
Slop Score (EQ-Bench + ours)5% weight
91.6
87.3
Visual Cue Quality4% weight
87.7
83.0

Everything else that differs

Kimi K3 (thinking)MiniMax M3
Overall / 10089.282.4
Writing Elo24161790
Run-to-run spread (± overall std)1.6505.610
Cost per script (USD)0.1870.019
Avg latency (s)267.2185.4
Open weightsYesYes

Full scorecards: Kimi K3 (thinking) · MiniMax M3. How scoring works: methodology.

← All comparisons