Towards AITowards AIToneBench

Every Thinking Machines model, side by side

Thinking Machines, Mira Murati's lab, has one model on the board so far, and it isn't a contender for scripts: Inkling lands around #115, well below the top tier. Its default and high-effort settings are both open weights, and they finish 5 Elo apart. This family page is short because the result is short.

Best writer
Inkling#115 · Elo 1319 · $0.035 per script
Cheapest
Inkling (high)#116 · Elo 1314 · $0.033 per script

All 2 configurations

Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.

Board rankConfigurationEloOverallToneCraftSubstanceHookYouTubeFlowSlopLengthCuesCost / scriptTime / script
#115Inkling · open1319 ±4078.378.680.880.083.174.876.281.770.380.5$0.0350.7 min
#116Inkling (high) · open1314 ±4477.977.980.781.081.474.476.082.068.080.7$0.0330.7 min

Which one should you use?

I'd run Inkling at its default and skip high effort. The two Elo ranges overlap, so the board can't separate them. They also cost about the same ($0.035 against $0.033), and both take around 0.7 minutes per script. The stat strip names Inkling (high) as the cheapest, but that's a sliver of a cent on measured cost, so don't read it as a recommendation.

Where I land: Inkling is fine if you already use it, but at #115 it's not what I'd pick for writing. For comparison, GLM-5.3 Flash sits at #18 for $0.015 per script, and that's where I'd look instead.

Thinking levels

Only Inkling ran at more than one reasoning-effort setting here. Its page shows what each step up buys.

Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.

← All comparisons