Towards AITowards AIToneBench

Claude Opus 5 (max) vs Kimi K3

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Opus 5 (max) leads overall, 90.8 to 89.4.

Claude Opus 5 (max)
#1 Elo 2649 · 90.8/100
Kimi K3
#8 Elo 2438 · 89.4/100
Cost / script
$0.120 vs $0.263
Human baseline
92.7 Claude Opus 5 (max) falls below it · Kimi K3 falls below it

The verdict

This is the closest an open-weights model has come to the top of this board. Opus 5 at max effort is still #1 with 2649.3 Elo against Kimi K3's 2437.5, and the confidence intervals don't touch, so the order is real. But look at the actual writing: 90.80 against 89.36. Under a point and a half. Kimi wins length adherence outright, 90.46 to 89.5, which on a 4,000 word script is not a small thing. Opus takes voice, craft, and hooks, and its cue quality is the gap I'd actually notice as an editor: 89.56 against 88.30. Then the part that changes the decision for a lot of teams. Opus costs about 12 cents a script, Kimi about 26. Yes, the open model is the expensive one here, mostly because it writes long and slow, 367 seconds against 244. So you're not picking Kimi to save money. You're picking it because you want open weights you can host yourself and scripts that land within a point and a half of the best closed model on the board. Both sit under the 92.74 human baseline, so neither is replacing the editorial pass yet.

Pick Opus 5 (max) if you want the best scripts on the board and the strongest visual cues, and 12 cents a script is not the constraint.
Pick Kimi K3 if you want open weights you can run yourself, tighter length discipline, and you can live with a point and a half and a slower draft.

Metric by metric

Blue bars: Claude Opus 5 (max). Orange bars: Kimi K3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
91.2
89.6
Writing Craft & Clarity13% weight
91.1
89.3
Substance, Accuracy & Value15% weight
90.6
89.5
Continuity & Emotion14% weight
90.4
88.3
YouTube Best Practices12% weight
89.6
88.4
Hook Strength10% weight
92.0
89.9
Length Adherence8% weight
89.7
90.5
Slop Score (EQ-Bench + ours)5% weight
93.3
92.0
Visual Cue Quality4% weight
89.6
88.3

Everything else that differs

Claude Opus 5 (max)Kimi K3
Overall / 10090.889.4
Writing Elo26492438
Run-to-run spread (± overall std)1.2901.570
Cost per script (USD)0.1200.263
Avg latency (s)243.8367.1
Open weightsNoYes

Full scorecards: Claude Opus 5 (max) · Kimi K3. How scoring works: methodology.

← All comparisons