Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 82.4.
Both open weights, and the price gap is the story: MiniMax M3 is under two cents a script, Kimi K3 is 26 cents. Thirteen times. For that you get #8 against #42 and 89.39 overall against 82.43. Seven and a half points. The widest single gap is YouTube structure, 88.39 to 77.73, and continuity is right behind at 88.28 to 79.16. Both describe the same failure: MiniMax loses the thread across a long script. Its hooks are fine at 86.12 and its voice is better than its rank suggests at 83.80, so the sentences are not the problem, the architecture is. MiniMax is also more than twice as fast, 185 seconds against 367. So the honest framing: if a human owns the outline and the model writes sections into it, MiniMax at two cents does that job. If the model owns the whole arc, the seven and a half points are exactly what you're paying Kimi for.
Pick Kimi K3 if the model has to carry the structure of the whole script by itself.
Pick MiniMax M3 if you own the outline and want open weights filling it in, thirteen times cheaper and twice as fast.
Blue bars: Kimi K3. Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 89.4 | 82.4 |
| Writing Elo | 2438 | 1790 |
| Run-to-run spread (± overall std) | 1.570 | 5.610 |
| Cost per script (USD) | 0.263 | 0.019 |
| Avg latency (s) | 367.1 | 185.4 |
| Open weights | Yes | Yes |
Full scorecards: Kimi K3 · MiniMax M3. How scoring works: methodology.
← All comparisons