Towards AITowards AIToneBench

Thinking levels · Thinking Machines

Does Inkling write better when it thinks harder?

Honestly, no. Inkling at its default and at high effort land 5 Elo apart, well inside each other's ranges, so the board can't tell them apart. The bill and the wait barely move either, which makes this an easy call.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
high1314 ±44-77.9$0.0330.7 min3,140#116
default (no effort flag)1319 ±40-78.3$0.0350.7 min3,636#115

The 95% Elo ranges overlap: Inkling runs from 1281 to 1362, and Inkling (high) from 1273 to 1360.

Which setting should you use?

Use the default and stop worrying about it. #115 against #116 is a rounding error, and you'd pay $0.035 or $0.033 per script either way. If your stack already asks for high effort, leave it. You're not losing anything. To be fair, it's a cheap, quick model. The question I'd actually ask is whether Inkling is the right model for scripts at all, since both settings sit well down the main board.

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every Thinking Machines model side by side.