Towards AITowards AIToneBench

Every xAI model, side by side

With xAI, which Grok you pick matters way more than how you set it. The 13 configurations here run from Grok 4.7 (high) at #26 on the full board all the way down to Grok 4 Fast at #154. And that's not a smooth ladder. The Grok 4.6 and Grok 4.7 settings bunch together at the top, then there's a cliff, and the very next model, Grok 4.5, falls to #90.

Best writer
Grok 4.7 (high)#26 · Elo 2166 · $0.311 per script
Best value (within 200 Elo of the best)
Grok 4.7 (low)#39 · Elo 2092 · $0.060 per script
Cheapest
Grok 3#150 · Elo 893 · $0.022 per script

All 13 configurations

Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.

Board rankConfigurationEloOverallToneCraftSubstanceHookYouTubeFlowSlopLengthCuesCost / scriptTime / script
#26Grok 4.7 (high)2166 ±4587.690.388.588.387.788.787.692.274.188.1$0.3118.7 min
#32Grok 4.72143 ±4687.690.489.388.888.488.287.692.471.487.7$0.2677.7 min
#39Grok 4.7 (low)2092 ±4486.990.289.788.887.987.588.090.865.485.5$0.0601.1 min
#41Grok 4.7 (medium)2086 ±5187.090.089.088.687.188.487.691.967.487.1$0.1904.7 min
#49Grok 4.62026 ±4586.487.686.686.879.587.185.590.888.087.6$0.2044.1 min
#50Grok 4.6 (high)1977 ±5186.087.786.486.479.086.685.790.785.687.6$0.2113.6 min
#90Grok 4.51590 ±4281.982.183.684.281.282.280.788.873.280.0$0.0630.9 min
#122Grok 4.20 (reasoning)1238 ±4977.277.579.277.681.076.075.784.076.360.3$0.0320.9 min
#142Grok 4.3 (high)1062 ±4974.172.477.078.578.073.473.987.257.268.0$0.0301.0 min
#145Grok 4 Fast (reasoning)1011 ±5073.672.077.078.077.272.873.687.656.763.5$0.0290.9 min
#149Grok 4.3895 ±4271.771.477.675.879.169.473.887.543.855.5$0.0220.5 min
#150Grok 3893 ±5971.370.777.576.478.069.073.686.343.154.1$0.0220.5 min
#154Grok 4 Fast830 ±4470.269.876.975.078.266.872.386.441.751.1$0.0220.4 min

Which one should you use?

If you want the best script Grok can write, that's Grok 4.7 (high), at $0.311 and 8.7 minutes per script. But for anything at volume, I'd start with Grok 4.7 (low). It's 74 Elo behind, close enough that their Elo ranges overlap. It also costs $0.060 and comes back in 1.1 minutes. The top setting asks for 5.2 times the money and a much longer wait for a gap the board can't fully confirm.

On paper the budget pick is Grok 3 at $0.022, and I wouldn't use it: Grok 4.7 (low) costs more, still just cents, and sits 1199 Elo higher. In this family, the cheap old models aren't a bargain.

Thinking levels, model by model

Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.

Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.

← All comparisons