Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 89.2.
Kimi K3 in thinking mode sits at #12, one place behind its own default config, which tells you most of what you need to know: the extra reasoning is not buying it much on writing tasks. Against Opus 5 at max effort the board reads 2649.3 to 2416.4 Elo, intervals clear of each other, and 90.80 to 89.25 on the score. Where thinking mode does earn its keep is length: 91.10 adherence, the best of any model in this comparison and nearly three points above Opus. On a 4,000 word brief that discipline shows. Everywhere else Opus leads by one to three points, and the widest gap is continuity and emotional flow, 90.5 against 88.28, which is the metric that decides whether a script reads like one person talking or a set of well-written paragraphs stapled together. Cost is 19 cents a script for Kimi against 12 for Opus. Open weights are the real argument here, not price.
Pick Opus 5 (max) if you want the strongest flow and voice on the board, and the best hooks by a clear margin.
Pick Kimi K3 (thinking) if you need open weights and the tightest length control here, and you're fine trading a point and a half of overall quality for it.
Blue bars: Claude Opus 5 (max). Orange bars: Kimi K3 (thinking). Same 0–100 scale; the bold bar wins that metric.
| Claude Opus 5 (max) | Kimi K3 (thinking) | |
|---|---|---|
| Overall / 100 | 90.8 | 89.2 |
| Writing Elo | 2649 | 2416 |
| Run-to-run spread (± overall std) | 1.290 | 1.650 |
| Cost per script (USD) | 0.120 | 0.187 |
| Avg latency (s) | 243.8 | 267.2 |
| Open weights | No | Yes |
Full scorecards: Claude Opus 5 (max) · Kimi K3 (thinking). How scoring works: methodology.
← All comparisons