Towards AITowards AIToneBench

Claude Fable 5 (max) vs Kimi K3 (thinking)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 89.2.

Claude Fable 5 (max)
#2 Elo 2575 · 90.3/100
Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
Cost / script
$0.908 vs $0.187
Human baseline
92.7 Claude Fable 5 (max) falls below it · Kimi K3 (thinking) falls below it

The verdict

Fable 5 wins this one on the board, but only barely on the page. Their Elo confidence intervals don't overlap, so the ranking is real: Fable second, Kimi K3 seventh. Then the scores themselves, 90.21 against 89.5, a gap of under a point. Head to head the judges keep picking Fable. Graded on their own, the two sets of scripts land in almost the same place. Where they actually differ: Fable handles numbers and slop better, 90.12 vs 87.24 on the anti-slop metric, and its visual cues are noticeably stronger. Kimi punches back on length adherence, 91.10 to Fable's 93.49, the one metric where it still comes out ahead. Then the part that decides it for a lot of people: Kimi is open weights and costs about 19 cents per script versus $0.98 for Fable. Five and a half times cheaper for a gap of three quarters of a point. Kimi is slower, over five minutes per script, but for drafting that rarely matters. Both still sit under the 92.74 human baseline, so my job is safe for now ;)

Pick Fable 5 if you want the cleanest, lowest-slop scripts with stronger visual cues and the budget isn't the constraint.
Pick Kimi K3 if you want writing that scores within a point of second place, open weights, and a roughly 5.5x lower bill per script.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: Kimi K3 (thinking). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
90.4
89.3
Writing Craft & Clarity13% weight
89.6
89.3
Substance, Accuracy & Value15% weight
90.2
89.2
Continuity & Emotion14% weight
89.1
88.3
YouTube Best Practices12% weight
89.0
87.8
Hook Strength10% weight
90.1
90.2
Length Adherence8% weight
93.5
91.1
Slop Score (EQ-Bench + ours)5% weight
93.7
91.6
Visual Cue Quality4% weight
90.0
87.7

Everything else that differs

Claude Fable 5 (max)Kimi K3 (thinking)
Overall / 10090.389.2
Writing Elo25752416
Run-to-run spread (± overall std)1.6101.650
Cost per script (USD)0.9080.187
Avg latency (s)452.8267.2
Open weightsNoYes

Full scorecards: Claude Fable 5 (max) · Kimi K3 (thinking). How scoring works: methodology.

← All comparisons