Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 82.4.
The headline says Kimi K3 at 89.2 against MiniMax M3 at 81.8, but the number that actually decides this matchup is the variance. MiniMax's overall spread is 5.61 points, nearly four times Kimi's, and its length adherence swings hardest of any metric here. In plain English: sometimes it lands the target, sometimes it's way off, and you won't know until you read it. On my learn-AI-engineering task it dropped to 75.5 while Kimi stayed above 88 on every task. That unreliability matters more to me than the average, because a model I have to babysit isn't saving me time. Credit to MiniMax though: it's open weights, the hooks are respectable at 86.1, and at roughly 2 cents per script, estimated, it's a fraction of Kimi's cost. The Elo intervals don't overlap, so the ranking itself isn't in doubt. If you want cheap open-weights drafts and you'll review every one, MiniMax is fine. If you want to trust what comes out, Kimi.
Pick Kimi K3 if you need output you can trust run after run, across every task type.
Pick MiniMax M3 if you want very cheap open-weights drafts and plan to review every single one.
Blue bars: Kimi K3 (thinking). Orange bars: MiniMax M3. Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 (thinking) | MiniMax M3 | |
|---|---|---|
| Overall / 100 | 89.2 | 82.4 |
| Writing Elo | 2416 | 1790 |
| Run-to-run spread (± overall std) | 1.650 | 5.610 |
| Cost per script (USD) | 0.187 | 0.019 |
| Avg latency (s) | 267.2 | 185.4 |
| Open weights | Yes | Yes |
Full scorecards: Kimi K3 (thinking) · MiniMax M3. How scoring works: methodology.
← All comparisons