Towards AITowards AIToneBench

Claude Fable 5 (max) vs Kimi K3

Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Claude Fable 5 (max) leads overall, 89.3 to 88.2.

Claude Fable 5 (max)
#1 Elo 2321 · 89.3/100
Kimi K3
#11 Elo 2180 · 88.2/100
Cost / script
$0.256 vs $0.260
Human baseline
90.2 Claude Fable 5 (max) falls below it · Kimi K3 falls below it

The verdict

Fable 5 wins this one, and the intervals say it cleanly: it sits at #1 with 2320.9 Elo against Kimi K3 at #11 with 2180.2, and the confidence intervals don't overlap, so the 140.7-point gap is real. Overall it's 89.3 to 88.19, a 1.11-point edge. Fable takes ten of the eleven metrics: length adherence by 3.13 points, 93.2 to 90.07, the numbers-and-slop metric by 3.36, 89.63 to 86.27, and continuity and emotion 88.26 to 86.8. Kimi's one win is hook strength, 90.19 to Fable's 89.39, and it punches harder head to head, ranking #2 pairwise where Fable sits #7. The records still lean Fable, 1185 wins and a single loss against Kimi's 1041 and 6. What's different here is the bill. Kimi is open weights, but at about twenty-six cents per script it costs essentially the same as Fable, so the usual cheap-open-model trade isn't on the table. What Kimi does offer is speed, turning a script around in less than half Fable's time. When the money is a wash, the better writer wins.

Pick Claude Fable 5 (max effort + 4.8 fallback) if you want the #1 board result, the cleanest length adherence, and the strongest handling of numbers and slop at essentially the same per-script cost as Kimi.
Pick Kimi K3 if you want open weights, the stronger hook at 90.19 to Fable's 89.39, and drafts in less than half the time, accepting a 1.11-point overall gap at the same price.

Metric by metric

Blue bars: Claude Fable 5 (max). Orange bars: Kimi K3. Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.4
88.3
Writing Craft & Clarity13% weight
88.8
88.5
Substance, Accuracy & Value15% weight
88.5
87.6
Continuity & Emotion14% weight
88.3
86.8
YouTube Best Practices12% weight
89.0
87.3
Hook Strength10% weight
89.4
90.2
Length Adherence8% weight
93.2
90.1
Slop Score (EQ-Bench + ours)5% weight
93.3
90.9
Visual Cue Quality4% weight
84.8
84.6

Everything else that differs

Claude Fable 5 (max)Kimi K3
Overall / 10089.388.2
Writing Elo23212180
Run-to-run spread (± overall std)1.5501.680
Cost per script (USD)0.2560.260
Avg latency (s)546.9234.3
Open weightsNoYes

Full scorecards: Claude Fable 5 (max) · Kimi K3. How scoring works: methodology.

← All comparisons