Towards AITowards AIToneBench

Thinking levels · xAI

Does Grok 4.6 write better when it thinks harder?

Not on this board. We ran Grok 4.6 at two settings, the effort flag on high and no flag at all, and asking for more thinking didn't buy a better script. The default setting comes out ahead by 49 Elo, which is inside the noise.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
high1977 ±51-86.0$0.2113.6 min11,633#50
default (no effort flag)2026 ±45-86.4$0.2044.1 min12,451#49

The 95% Elo ranges overlap: Grok 4.6 runs from 1985 to 2074, and Grok 4.6 (high) from 1924 to 2026.

Which setting should you use?

Leave the flag off. Both settings cost about the same per script, $0.204 without the flag and $0.211 with it, and the wait is similar. When extra effort doesn't move the score, there's no reason to ask for it. If you're choosing a Grok for scripts right now, though, look at Grok 4.7 (low) first: it scores at least as well, and Grok 4.6 costs 3.4 times as much per script.

Other xAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every xAI model side by side.