Towards AITowards AIToneBench

Thinking levels · xAI

Does Grok 4.7 write better when it thinks harder?

A little, yes, but you pay a lot for that little. Going from low to high adds 74 Elo, while the script costs 5.2 times more and takes 8.7 minutes instead of 1.1. And the step in between barely moves: medium lands right next to low.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
low2092 ±44-86.9$0.0601.1 min1,118#39
medium2086 ±51-787.0$0.1904.7 min18,253#41
high2166 ±45+8187.6$0.3118.7 min38,750#26
default (no effort flag)2143 ±46-87.6$0.2677.7 min34,196#32

The 95% Elo ranges overlap: Grok 4.7 (low) runs from 2044 to 2132, and Grok 4.7 (high) from 2121 to 2211.

Which setting should you use?

I'd use low for most work. It costs $0.060, and high would only beat it in about 60% of script comparisons, with a tie counting as half a win. If one script really matters, high is the one to try, as long as you know that gap could still be noise. Skip medium, though: it costs 3.2 times as much as low for no visible gain. Leaving the flag off isn't a shortcut either. It scores close to high and costs nearly as much, at $0.267.

Other xAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every xAI model side by side.