Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 leads overall, 89.4 to 88.4.
Best open-weights model against GPT's top rung, and the intervals are clear of each other this time, so the order stands: Kimi K3 #8 at 89.39, GPT-5.6 Sol ultra #16 at 88.40. The split is the same one that runs through every Kimi-versus-GPT page on this board. Kimi wins the writing: hooks 89.88 against 86.39, tone 89.61 against 86.99, continuity 88.28 against 85.72. GPT wins the mechanics: length adherence 94.34 against 90.46, cue quality 91.17 against 88.30, number discipline 90.19 against 87.36. One honest caveat, as always: GPT-5.6 Sol is one of my three judges and rates its own family warmly, though the other two seats agree on the direction here. Ultra costs about 14 cents a script against Kimi's 26 and is slower, seven minutes against six. I keep landing the same place: GPT hands me tidier drafts, Kimi hands me better writing, and Kimi being open weights breaks the tie for me.
Pick Kimi K3 for the better read: it wins hooks, tone and continuity, and the weights are yours to run.
Pick GPT-5.6 Sol ultra when mechanical discipline matters most: best-in-class length control and visual cues, at half Kimi's price.
Blue bars: Kimi K3. Orange bars: GPT-5.6 Sol (ultra). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 | GPT-5.6 Sol (ultra) | |
|---|---|---|
| Overall / 100 | 89.4 | 88.4 |
| Writing Elo | 2438 | 2316 |
| Run-to-run spread (± overall std) | 1.570 | 1.970 |
| Cost per script (USD) | 0.263 | 0.137 |
| Avg latency (s) | 367.1 | 420.9 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 · GPT-5.6 Sol (ultra). How scoring works: methodology.
← All comparisons