Towards AITowards AIToneBench

Kimi K3 vs MiniMax M3

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 82.4.

Kimi K3
#8 Elo 2438 · 89.4/100
MiniMax M3
#42 Elo 1790 · 82.4/100
Cost / script
$0.263 vs $0.019
Human baseline
92.7 Kimi K3 falls below it · MiniMax M3 falls below it

The verdict

Both open weights, and the price gap is the story: MiniMax M3 is under two cents a script, Kimi K3 is 26 cents. Thirteen times. For that you get #8 against #42 and 89.39 overall against 82.43. Seven and a half points. The widest single gap is YouTube structure, 88.39 to 77.73, and continuity is right behind at 88.28 to 79.16. Both describe the same failure: MiniMax loses the thread across a long script. Its hooks are fine at 86.12 and its voice is better than its rank suggests at 83.80, so the sentences are not the problem, the architecture is. MiniMax is also more than twice as fast, 185 seconds against 367. So the honest framing: if a human owns the outline and the model writes sections into it, MiniMax at two cents does that job. If the model owns the whole arc, the seven and a half points are exactly what you're paying Kimi for.

Pick Kimi K3 if the model has to carry the structure of the whole script by itself.
Pick MiniMax M3 if you own the outline and want open weights filling it in, thirteen times cheaper and twice as fast.

Metric by metric

Blue bars: Kimi K3. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.6
83.8
Writing Craft & Clarity13% weight
89.3
83.9
Substance, Accuracy & Value15% weight
89.5
81.8
Continuity & Emotion14% weight
88.3
79.2
YouTube Best Practices12% weight
88.4
77.7
Hook Strength10% weight
89.9
86.1
Length Adherence8% weight
90.5
82.6
Slop Score (EQ-Bench + ours)5% weight
92.0
87.3
Visual Cue Quality4% weight
88.3
83.0

Everything else that differs

Kimi K3MiniMax M3
Overall / 10089.482.4
Writing Elo24381790
Run-to-run spread (± overall std)1.5705.610
Cost per script (USD)0.2630.019
Avg latency (s)367.1185.4
Open weightsYesYes

Full scorecards: Kimi K3 · MiniMax M3. How scoring works: methodology.

← All comparisons