Towards AITowards AIToneBench

Thinking levels · DeepSeek

Does DeepSeek V4 Pro write better when it thinks harder?

Not in a way this board can confirm. Setting DeepSeek V4 Pro to xhigh lands 44 Elo above the default, but the two ranges overlap. Meanwhile you pay 1.5 times as much per script and wait 1.8 minutes instead of 1.0.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
xhigh1370 ±56-79.0$0.0441.8 min7,080#111
default (no effort flag)1326 ±39-78.7$0.0301.0 min2,284#113

The 95% Elo ranges overlap: DeepSeek V4 Pro (default) runs from 1287 to 1366, and DeepSeek V4 Pro (xhigh) from 1315 to 1427.

Which setting should you use?

Leave it on default. The extra thinking costs more and takes longer, all for a gap that sits inside the noise. That's a weird result next to DeepSeek V4 Pro 0813, where max effort is worth 476 Elo, a huge jump. This version just doesn't seem to get much out of xhigh. And to be honest, I wouldn't pick it for scripts at either setting: DeepSeek V4.1 Flash (max) ranks #27 for $0.013 a script.

Other DeepSeek ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every DeepSeek model side by side.