Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 89.4.
This is the closest an open-weights model has come to the top of this board. Opus 5 at max effort is still #1 with 2649.3 Elo against Kimi K3's 2437.5, and the confidence intervals don't touch, so the order is real. But look at the actual writing: 90.80 against 89.36. Under a point and a half. Kimi wins length adherence outright, 90.46 to 89.5, which on a 4,000 word script is not a small thing. Opus takes voice, craft, and hooks, and its cue quality is the gap I'd actually notice as an editor: 89.56 against 88.30. Then the part that changes the decision for a lot of teams. Opus costs about 12 cents a script, Kimi about 26. Yes, the open model is the expensive one here, mostly because it writes long and slow, 367 seconds against 244. So you're not picking Kimi to save money. You're picking it because you want open weights you can host yourself and scripts that land within a point and a half of the best closed model on the board. Both sit under the 92.74 human baseline, so neither is replacing the editorial pass yet.
Pick Opus 5 (max) if you want the best scripts on the board and the strongest visual cues, and 12 cents a script is not the constraint.
Pick Kimi K3 if you want open weights you can run yourself, tighter length discipline, and you can live with a point and a half and a slower draft.
Blue bars: Claude Opus 5 (max). Orange bars: Kimi K3. Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | Kimi K3 | |
|---|---|---|
| Overall / 100 | 90.8 | 89.4 |
| Writing Elo | 2649 | 2438 |
| Run-to-run spread (± overall std) | 1.290 | 1.570 |
| Cost per script (USD) | 0.120 | 0.263 |
| Avg latency (s) | 243.8 | 367.1 |
| Open weights | No | Yes |
Full scorecards: Claude Opus 5 (max) · Kimi K3. How scoring works: methodology.
← All comparisons