Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 88.2.
Fable 5 wins this one, and the intervals say it cleanly: it sits at #1 with 2320.9 Elo against Kimi K3 at #11 with 2180.2, and the confidence intervals don't overlap, so the 140.7-point gap is real. Overall it's 89.3 to 88.19, a 1.11-point edge. Fable takes ten of the eleven metrics: length adherence by 3.13 points, 93.2 to 90.07, the numbers-and-slop metric by 3.36, 89.63 to 86.27, and continuity and emotion 88.26 to 86.8. Kimi's one win is hook strength, 90.19 to Fable's 89.39, and it punches harder head to head, ranking #2 pairwise where Fable sits #7. The records still lean Fable, 1185 wins and a single loss against Kimi's 1041 and 6. What's different here is the bill. Kimi is open weights, but at about twenty-six cents per script it costs essentially the same as Fable, so the usual cheap-open-model trade isn't on the table. What Kimi does offer is speed, turning a script around in less than half Fable's time. When the money is a wash, the better writer wins.
Pick Claude Fable 5 (max effort + 4.8 fallback) if you want the #1 board result, the cleanest length adherence, and the strongest handling of numbers and slop at essentially the same per-script cost as Kimi.
Pick Kimi K3 if you want open weights, the stronger hook at 90.19 to Fable's 89.39, and drafts in less than half the time, accepting a 1.11-point overall gap at the same price.
Blue bars: Claude Fable 5 (max). Orange bars: Kimi K3. Same 0–100 scale; the bold bar wins that metric.
| Claude Fable 5 (max) | Kimi K3 | |
|---|---|---|
| Overall / 100 | 89.3 | 88.2 |
| Writing Elo | 2321 | 2180 |
| Run-to-run spread (± overall std) | 1.550 | 1.680 |
| Cost per script (USD) | 0.256 | 0.260 |
| Avg latency (s) | 546.9 | 234.3 |
| Open weights | No | Yes |
Full scorecards: Claude Fable 5 (max) · Kimi K3. How scoring works: methodology.
← All comparisons