Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 88.4. Their confidence intervals overlap, so treat the order as close rather than settled.
This one is close, and both sit near the top of the board. Kimi K3 outranks GPT-5.6 Sol on Elo, 2416 to 2315, though the confidence intervals brush at the edges, so call it a strong lean rather than settled. Either way, they win differently. Kimi writes the better openings, 90.2 on hooks against 87.0, and lands closer to my actual voice, 89.3 versus 87.0 on tone. GPT-5.6 is the tidier craftsman: stronger visual cues at 91.0 versus 87.8 and better number discipline at 90.6 versus 87.2, so fewer stat dumps I'd have to rewrite. One honest caveat: GPT-5.6 Sol is also one of my three judges, and it rates its own scripts higher than the other two do. On practicality, GPT-5.6 runs about 16 cents per script against Kimi's 20 and finishes in under a third of the time. For me, Kimi being open weights matters more than either number. And both still trail my human baseline of 93.0, which honestly, I find reassuring.
Pick Kimi K3 if you want the stronger hooks and closer voice match, plus open weights, and you can live with slower generations.
Pick GPT-5.6 Sol if you want cleaner visual cues, tighter number discipline, and a faster, cheaper closed option.
Blue bars: Kimi K3 (thinking). Orange bars: GPT-5.6 Sol (high). Same 0–100 scale; the bold bar wins that metric.
| Kimi K3 (thinking) | GPT-5.6 Sol (high) | |
|---|---|---|
| Overall / 100 | 89.2 | 88.4 |
| Writing Elo | 2416 | 2315 |
| Run-to-run spread (± overall std) | 1.650 | 1.750 |
| Cost per script (USD) | 0.187 | 0.163 |
| Avg latency (s) | 267.2 | 106.4 |
| Open weights | Yes | No |
Full scorecards: Kimi K3 (thinking) · GPT-5.6 Sol (high). How scoring works: methodology.
← All comparisons