Towards AITowards AIToneBench

Kimi K3 (thinking) vs GPT-5.6 Sol (high)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Kimi K3 (thinking) leads overall, 89.2 to 88.4. Their confidence intervals overlap, so treat the order as close rather than settled.

Kimi K3 (thinking)
#12 Elo 2416 · 89.2/100
GPT-5.6 Sol (high)
#17 Elo 2315 · 88.4/100
Cost / script
$0.187 vs $0.163
Human baseline
92.7 Kimi K3 (thinking) falls below it · GPT-5.6 Sol (high) falls below it

The verdict

This one is close, and both sit near the top of the board. Kimi K3 outranks GPT-5.6 Sol on Elo, 2416 to 2315, though the confidence intervals brush at the edges, so call it a strong lean rather than settled. Either way, they win differently. Kimi writes the better openings, 90.2 on hooks against 87.0, and lands closer to my actual voice, 89.3 versus 87.0 on tone. GPT-5.6 is the tidier craftsman: stronger visual cues at 91.0 versus 87.8 and better number discipline at 90.6 versus 87.2, so fewer stat dumps I'd have to rewrite. One honest caveat: GPT-5.6 Sol is also one of my three judges, and it rates its own scripts higher than the other two do. On practicality, GPT-5.6 runs about 16 cents per script against Kimi's 20 and finishes in under a third of the time. For me, Kimi being open weights matters more than either number. And both still trail my human baseline of 93.0, which honestly, I find reassuring.

Pick Kimi K3 if you want the stronger hooks and closer voice match, plus open weights, and you can live with slower generations.
Pick GPT-5.6 Sol if you want cleaner visual cues, tighter number discipline, and a faster, cheaper closed option.

Metric by metric

Blue bars: Kimi K3 (thinking). Orange bars: GPT-5.6 Sol (high). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
89.3
87.3
Writing Craft & Clarity13% weight
89.3
88.1
Substance, Accuracy & Value15% weight
89.2
90.1
Continuity & Emotion14% weight
88.3
86.1
YouTube Best Practices12% weight
87.8
86.5
Hook Strength10% weight
90.2
87.5
Length Adherence8% weight
91.1
91.7
Slop Score (EQ-Bench + ours)5% weight
91.6
93.4
Visual Cue Quality4% weight
87.7
91.0

Everything else that differs

Kimi K3 (thinking)GPT-5.6 Sol (high)
Overall / 10089.288.4
Writing Elo24162315
Run-to-run spread (± overall std)1.6501.750
Cost per script (USD)0.1870.163
Avg latency (s)267.2106.4
Open weightsYesNo

Full scorecards: Kimi K3 (thinking) · GPT-5.6 Sol (high). How scoring works: methodology.

← All comparisons