Towards AITowards AIToneBench

Thinking levels · DeepSeek

Does DeepSeek V4 Pro 0813 write better when it thinks harder?

This one really does. We ran DeepSeek V4 Pro 0813 at two settings, the provider default and max, and turning it up adds 476 Elo. You pay for that twice: 2.2 times the cost per script, and 3.2 minutes of waiting instead of 1.1.

X axis

Both settings we ran

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
max1884 ±68-84.6$0.0443.2 min15,642#68
default (no effort flag)1408 ±61-79.2$0.0201.1 min4,152#108

Which setting should you use?

If you're on this model, use max. The default run ranks #108, max pulls it up to #68, and the two ranges don't overlap, so that's strong evidence of a real gain. The catch is what sits above it. Even at max, it costs 3.5 times what DeepSeek V4.1 Flash (max) charges and still ranks below it, so for scripts I'd reach for the cheaper Flash model first.

Other DeepSeek ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every DeepSeek model side by side.