Towards AITowards AIToneBench

Claude Opus 5 (max) vs Kimi K3 (thinking)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 89.2.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
Cost / script
$0.120 vs $0.187
Human baseline
92.7 Claude Opus 5 (max) falls below it · Kimi K3 (thinking) falls below it

The verdict

Kimi K3 in thinking mode sits at #12, one place behind its own default config, which tells you most of what you need to know: the extra reasoning is not buying it much on writing tasks. Against Opus 5 at max effort the board reads 2649.3 to 2416.4 Elo, intervals clear of each other, and 90.80 to 89.25 on the score. Where thinking mode does earn its keep is length: 91.10 adherence, the best of any model in this comparison and nearly three points above Opus. On a 4,000 word brief that discipline shows. Everywhere else Opus leads by one to three points, and the widest gap is continuity and emotional flow, 90.5 against 88.28, which is the metric that decides whether a script reads like one person talking or a set of well-written paragraphs stapled together. Cost is 19 cents a script for Kimi against 12 for Opus. Open weights are the real argument here, not price.

Pick Opus 5 (max) if you want the strongest flow and voice on the board, and the best hooks by a clear margin.
Pick Kimi K3 (thinking) if you need open weights and the tightest length control here, and you're fine trading a point and a half of overall quality for it.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: Kimi K3 (thinking). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
89.3
Writing Craft & Clarity13% weight
91.1
89.3
Substance, Accuracy & Value15% weight
90.6
89.2
Continuity & Emotion14% weight
90.4
88.3
YouTube Best Practices12% weight
89.6
87.8
Hook Strength10% weight
92.0
90.2
Length Adherence8% weight
89.7
91.1
Slop Score (EQ-Bench + ours)5% weight
93.3
91.6
Visual Cue Quality4% weight
89.6
87.7

Everything else that differs

Claude Opus 5 (max)Kimi K3 (thinking)
Overall / 10090.889.2
Writing Elo26492416
Run-to-run spread (± overall std)1.2901.650
Cost per script (USD)0.1200.187
Avg latency (s)243.8267.2
Open weightsNoYes

Full scorecards: Claude Opus 5 (max) · Kimi K3 (thinking). How scoring works: methodology.

← All comparisons