Towards AITowards AIToneBench

Every DeepSeek model, side by side

If you write a lot of scripts and the bill matters, start with DeepSeek. Every configuration here is open weights, all 15 of them, and none costs more than a few cents per script. The trap is the spread: 1015 Elo separates its best run from its worst. So choosing the right DeepSeek matters a lot more than choosing DeepSeek.

Best writer
DeepSeek V4.1 Flash (max)#27 · Elo 2165 · $0.013 per script
Best value (within 200 Elo of the best)
DeepSeek V4.1 Flash (default)#44 · Elo 2068 · $0.0077 per script
Cheapest
DeepSeek V4 Flash (default)#135 · Elo 1150 · $0.0014 per script

All 15 configurations

Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.

Board rankConfigurationEloOverallToneCraftSubstanceHookYouTubeFlowSlopLengthCuesCost / scriptTime / script
#27DeepSeek V4.1 Flash (max) · open2165 ±3687.887.888.986.387.687.186.690.492.084.0$0.0131.3 min
#44DeepSeek V4.1 Flash (default) · open2068 ±4786.887.188.686.087.385.085.689.189.682.2$0.00770.7 min
#45DeepSeek V4.1 Flash (high) · open2066 ±4486.887.088.786.187.885.786.289.887.679.9$0.00910.8 min
#60DeepSeek V4.1 Flash (low) · open1937 ±5285.586.287.985.187.183.585.189.184.078.0$0.00540.5 min
#68DeepSeek V4 Pro 0813 (max) · open1884 ±6884.684.885.784.486.084.082.187.786.181.8$0.0443.2 min
#71DeepSeek V4 Flash 0731 · open1859 ±7884.383.985.784.485.583.282.387.286.182.2$0.00783.1 min
#76DeepSeek V4.1 Flash (no thinking) · open1710 ±5283.584.386.384.385.880.683.087.378.877.6$0.00330.2 min
#99DeepSeek V4 Flash (native chat alias) · open1495 ±6080.680.683.481.682.574.078.884.183.578.2$0.00240.6 min
#108DeepSeek V4 Pro 0813 (default) · open1408 ±6179.278.781.780.383.577.977.084.572.178.2$0.0201.1 min
#111DeepSeek V4 Pro (xhigh) · open1370 ±5679.080.081.479.582.575.277.580.278.570.8$0.0441.8 min
#113DeepSeek V4 Pro (default) · open1326 ±3978.779.381.579.883.275.678.778.774.370.5$0.0301.0 min
#127DeepSeek V4 Flash (xhigh) · open1199 ±4677.077.379.376.881.374.975.182.270.876.1$0.00333.0 min
#132DeepSeek V4 Flash (native reasoner alias) · open1173 ±4575.975.479.978.280.869.475.181.466.378.8$0.00160.7 min
#134DeepSeek V3.2 · open1157 ±3876.274.780.478.080.774.076.576.367.575.0$0.00401.1 min
#135DeepSeek V4 Flash (default) · open1150 ±4075.875.179.777.781.369.575.281.865.679.1$0.00140.7 min

Which one should you use?

I'd use DeepSeek V4.1 Flash (max). It's #27 on the whole board at $0.013 a script, which is cheap enough that most people can stop there. If you're generating at real volume, DeepSeek V4.1 Flash (default) is the value pick: 97 Elo behind, for a cheaper and faster script at $0.0077.

Here's the odd part. Pro is the bigger, pricier tier, and it doesn't pay off for scripts: its best run, DeepSeek V4 Pro 0813 (max), costs 3.5 times as much as DeepSeek V4.1 Flash (max) and only reaches #68. And don't bother with the bottom of the price list. DeepSeek V4 Flash (default) costs $0.0014, but it ranks #135, and when every option is this cheap, saving there makes no sense.

Thinking levels, model by model

Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.

Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.

← All comparisons