Towards AITowards AIToneBench

Kimi K3

Moonshot (Kimi) · open weights · writing benchmark

#21 of 167 Elo 2217 Overall 88.5 Open weights
Rank
#21 of 167
Writing Elo
2217 ±35
Overall
88.5 / 100
Cost / script
$0.26 per script
Family
Moonshot (Kimi) 1st of 6
Type
Open 3rd of 53
Consistency
± 1.3 steadier than most
Avg tokens
26.5k in+out
Time / script
3.9 min avg, retries included
Head-to-head:vs Claude Opus 5.5 (max) · vs GLM-5.3 · vs GPT-5.6 Sol (ultra) · vs Grok 4.7 (high) · vs DeepSeek V4.1 Flash (max) · vs Qwen3.8 Flash · vs MiniMax M3

The short version

Kimi K3 lands in the upper third of ToneBench at #21 of 167, with a writing Elo of 2217 and an overall score of 88.5 out of 100. Those numbers come from it writing its own version of the scripts behind 10 of my What's AI videos, five times each, with every draft scored blind against mine (how the scoring works).

For scale, my own scripts score 90.9 through the same judges and rubric. That's the real bar here, not a literal 100. It sits 2.4 points under it.

Its Elo comes from comparing its average on each script with every other model's average on the same script, with gaps inside the run-to-run noise counted as draws. The 95% range is 2179 to 2248. That range overlaps 21 other configurations, ranked #11 to #32, so its exact spot inside that band is too close to call on this board. What the ranges can and can't tell you.

Its biggest edge is Length: 90.1, 16th on the board and about 19 points above the board average. Even its weakest, Slop (higher means cleaner) (91.2), sits above the board average, so there's no real weak spot to plan around.

It held steady across all 10 scripts, with its best on the career guide / hiring analysis (89.9).

It's also unusually consistent: its overall moves only ± 1.3 between runs of the same script, against ± 2.7 for the typical model. What you get on the first try is close to what you get every time.

It costs about $0.26 per script. GLM-5.3 Flash scores higher and costs about $0.007 per script, so on value this isn't the pick. Still, a few pricier configs score lower than it does.

It's the best Moonshot (Kimi) writer on the board, ahead of the other 5 Moonshot (Kimi) configs. If you're staying with Moonshot (Kimi), start here.

Among open-weights models, it's 3rd of 53. That's a strong showing for open weights, because voice and tone is usually where the closed frontier still pulls ahead.

Skill profile

Read the shape: the further a corner reaches, the stronger the model is on that metric. The dashed line is the board average and the dotted line is the board's best score on each metric, so a corner outside the dashed line beats the average there. Hover a point for the exact numbers.

50Tone?Craft?Substance?Flow?YouTube?Hook?Length?Slop?Cues?
Kimi K3Board averageBoard best per metricMax possible (100)

Metric by metric

Nine writing metrics, each scored 0–100 and blended into the overall by the editorial weight shown next to it. Rank is out of all 167 current-ranked models. The thick bar is this model; the thin line above it is the field's best on that metric (darker) and the one below is the field's average (lighter), on the same scale. The small ± is how much the score moves from run to run, and the ? next to each name says what the metric rewards.

Tone ?19% weight · +9.8 vs avg
best · Claude Opus 5.5 (max effort) · 91.088.4± 1.1 · 25thboard avg · 78.6
Craft ?13% weight · +8.5 vs avg
best · Claude Opus 5.5 (max effort) · 91.589.0± 1.0 · 23rdboard avg · 80.5
Substance ?15% weight · +8.2 vs avg
best · Claude Opus 5.5 (xhigh) · 90.787.6± 1.1 · 32ndboard avg · 79.4
Flow ?14% weight · +10.6 vs avg
best · Claude Opus 5.5 (max effort) · 90.987.6± 1.2 · 28thboard avg · 77.0
YouTube ?12% weight · +11.7 vs avg
best · Claude Opus 5.5 (max effort) · 91.788.4± 1.4 · 21stboard avg · 76.7
Hook ?10% weight · +7.7 vs avg
best · Claude Opus 5.5 (max effort) · 91.289.2± 1.2 · 14thboard avg · 81.6
Length ?8% weight · +19.3 vs avg
best · Claude Opus 5.5 (max effort) · 97.690.1± 7.1 · 16thboard avg · 70.8
Slop ?5% weight · +7.1 vs avg
best · Claude Opus 5.5 (max effort) · 94.991.2± 1.5 · 45thboard avg · 84.2
Cues ?4% weight · +8.2 vs avg
best · GPT-5.6 Sol (ultra) · 90.185.1± 2.7 · 59thboard avg · 77.0

Script by script

The same 10 real scripts every current-ranked model writes, each scored on its own as the mean of 5 runs. Different formats stress different skills: a model can nail a tight explainer and still stumble on a personal story, so look for the format closest to what you write.

Script 1
opinion / warning explainer
88.7
out of 100
Script 2
news-analysis / skeptical explainer
88.2
out of 100
Script 3
personal roadmap / opinion
87.9
out of 100
Script 4
short explainer
87.8
out of 100
Script 5
news-analysis / opinion explainer
88.7
out of 100
Script 6
founder announcement / personal origin story
89.2
out of 100
Script 7
personal technical walkthrough / agentic workflow case study
89.2
out of 100
Script 8
product release announcement / personal observation
87.6
out of 100
Script 9
career guide / hiring analysis
89.9
out of 100
Script 10
engineering process walkthrough / presentation adaptation
87.6
out of 100

Consistency: how much it moves between runs

A model whose drafts swing from run to run is harder to use than its average suggests. Every score on this page is the mean of 5 runs per script, and the ± is the run-to-run standard deviation.

Its steadiest is YouTube: ± 1.4 run to run, when the typical model swings ± 2.9.

MetricThis modelBoard medianVerdict
Tone ?88.4 ± 1.1± 1.7typical
Craft ?89.0 ± 1.0± 1.5typical
Substance ?87.6 ± 1.1± 2.0steadier than most
Flow ?87.6 ± 1.2± 1.8typical
YouTube ?88.4 ± 1.4± 2.9steadier than most
Hook ?89.2 ± 1.2± 2.0typical
Length ?90.1 ± 7.1± 8.6typical
Slop ?91.2 ± 1.5± 2.1typical
Cues ?85.1 ± 2.7± 3.2typical

Cost, time and tokens

Logged on every run. None of it touches the writing score, but it's often what decides whether a model fits your pipeline at all. Time is the average time to get one accepted script, retries included, and cost is priced per script at list price.

Time per script
3.9 minslower than the board median of 1.2 min
Prompt tokens in
11.5kstyle guide + brief + research packet
Tokens out
15.0kwell above the board median (a lot of thinking)
Cost per script
$0.260 ± 0.092measured from actual billed tokens
List price used
$3 / $15 per M tokinput / output

What each judge scored it

One judge can have taste of its own, which is why there are three, from three model families, each scoring the same 50 stored drafts blind with the identical rubric. The overall of 88.5 combines their scores with the parts computed in code (length, and half of the slop score). The highest and lowest judge are 4.0 points apart on its overall. That's tighter than the board median of 6.3, so the panel mostly agrees on this one. How the panel works.

Claude Opus 5
Anthropic
86.9
GPT-5.6 Sol (medium)
OpenAI
90.9
DeepSeek V4.1 Flash
DeepSeek
87.6

How we ran Kimi K3

Frequently asked questions

How good is Kimi K3 at writing?

It's #21 of 167 on ToneBench, with a writing Elo of 2217 and an overall score of 88.5/100. It wrote its own version of the scripts behind 10 of my What's AI videos 5 times each, and three judges from three model families scored every draft blind against mine. Length and half of the slop score are computed in code.

Is Kimi K3 the best Moonshot (Kimi) model for writing?

Yes. Of the 6 Moonshot (Kimi) configs on the board, Kimi K3 writes best.

Is Kimi K3 good value for the money?

It costs about $0.26 per script. GLM-5.3 Flash scores higher and costs less, so it isn't the value pick.

What are Kimi K3's strengths and weaknesses?

Against the board average, its biggest edge is Length (90.1, +19.3) and even its weakest, Slop (higher means cleaner) (91.2, +7.1), is above average. The nine-metric breakdown on this page shows the board average and board best for each.

How was Kimi K3 evaluated?

Through OpenRouter, using the exact model/route id moonshotai/kimi-k3, first run on 2026-09-03. Each script ran 5 times with provider-default sampling and no fine-tuning. The 'How we ran it' section and the methodology page have the rest.

Is Kimi K3 open source?

Yes, it's an open-weights model. Among open-weights models it's 3rd of 53 for writing.

← Back to the full ToneBench leaderboard