Towards AITowards AIToneBench

Thinking levels · DeepSeek

Does DeepSeek V4.1 Flash write better when it thinks harder?

Yes, and by a lot. Going from no thinking at all to max adds 455 Elo, enough that the max setting would win about 93% of its script comparisons against the no-thinking run. The usual reason to hold back doesn't really apply here: even at max, a script costs $0.013 and takes about 1.3 minutes.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptReasoning tokensBoard rank
no reasoning1710 ±52-83.5$0.00330.2 min-#76
low1937 ±52+22785.5$0.00540.5 min3,339#60
high2066 ±44+12886.8$0.00910.8 min9,409#45
max2165 ±36+9987.8$0.0131.3 min15,361#27
default (no effort flag)2068 ±47-86.8$0.00770.7 min7,210#44

One disclosure: DeepSeek V4.1 Flash itself also holds a seat on the judge panel. The panel spans three LLM families plus Jev, so no family scores itself alone, and the methodology explains how that works.

Which setting should you use?

I'd run it at max. It's the top-scoring setting, its range doesn't overlap the default one, and even at 3.8 times the no-thinking price, the bill stays tiny. If you want it faster, default is the pick: 97 Elo behind max, 0.7 minutes a script, and going up to high doesn't buy you anything the board can measure. What I wouldn't do is turn thinking off, because that's the biggest drop on the whole ladder.

Other DeepSeek ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every DeepSeek model side by side.