Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 90.3 to 89.2.
Fable 5 wins this one on the board, but only barely on the page. Their Elo confidence intervals don't overlap, so the ranking is real: Fable second, Kimi K3 seventh. Then the scores themselves, 90.21 against 89.5, a gap of under a point. Head to head the judges keep picking Fable. Graded on their own, the two sets of scripts land in almost the same place. Where they actually differ: Fable handles numbers and slop better, 90.12 vs 87.24 on the anti-slop metric, and its visual cues are noticeably stronger. Kimi punches back on length adherence, 91.10 to Fable's 93.49, the one metric where it still comes out ahead. Then the part that decides it for a lot of people: Kimi is open weights and costs about 19 cents per script versus $0.98 for Fable. Five and a half times cheaper for a gap of three quarters of a point. Kimi is slower, over five minutes per script, but for drafting that rarely matters. Both still sit under the 92.74 human baseline, so my job is safe for now ;)
Pick Fable 5 if you want the cleanest, lowest-slop scripts with stronger visual cues and the budget isn't the constraint.
Pick Kimi K3 if you want writing that scores within a point of second place, open weights, and a roughly 5.5x lower bill per script.
Blue bars: Claude Fable 5 (max). Orange bars: Kimi K3 (thinking). Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | Kimi K3 (thinking) | |
|---|---|---|
| Overall / 100 | 90.3 | 89.2 |
| Writing Elo | 2575 | 2416 |
| Run-to-run spread (± overall std) | 1.610 | 1.650 |
| Cost per script (USD) | 0.908 | 0.187 |
| Avg latency (s) | 452.8 | 267.2 |
| Open weights | No | Yes |
Full scorecards: Claude Fable 5 (max) · Kimi K3 (thinking). How scoring works: methodology.
← All comparisons