If you write a lot of scripts and the bill matters, start with DeepSeek. Every configuration here is open weights, all 15 of them, and none costs more than a few cents per script. The trap is the spread: 1015 Elo separates its best run from its worst. So choosing the right DeepSeek matters a lot more than choosing DeepSeek.
Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.
| Board rank | Configuration | Elo | Overall | Tone | Craft | Substance | Hook | YouTube | Flow | Slop | Length | Cues | Cost / script | Time / script |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| #27 | DeepSeek V4.1 Flash (max) · open | 2165 ±36 | 87.8 | 87.8 | 88.9 | 86.3 | 87.6 | 87.1 | 86.6 | 90.4 | 92.0 | 84.0 | $0.013 | 1.3 min |
| #44 | DeepSeek V4.1 Flash (default) · open | 2068 ±47 | 86.8 | 87.1 | 88.6 | 86.0 | 87.3 | 85.0 | 85.6 | 89.1 | 89.6 | 82.2 | $0.0077 | 0.7 min |
| #45 | DeepSeek V4.1 Flash (high) · open | 2066 ±44 | 86.8 | 87.0 | 88.7 | 86.1 | 87.8 | 85.7 | 86.2 | 89.8 | 87.6 | 79.9 | $0.0091 | 0.8 min |
| #60 | DeepSeek V4.1 Flash (low) · open | 1937 ±52 | 85.5 | 86.2 | 87.9 | 85.1 | 87.1 | 83.5 | 85.1 | 89.1 | 84.0 | 78.0 | $0.0054 | 0.5 min |
| #68 | DeepSeek V4 Pro 0813 (max) · open | 1884 ±68 | 84.6 | 84.8 | 85.7 | 84.4 | 86.0 | 84.0 | 82.1 | 87.7 | 86.1 | 81.8 | $0.044 | 3.2 min |
| #71 | DeepSeek V4 Flash 0731 · open | 1859 ±78 | 84.3 | 83.9 | 85.7 | 84.4 | 85.5 | 83.2 | 82.3 | 87.2 | 86.1 | 82.2 | $0.0078 | 3.1 min |
| #76 | DeepSeek V4.1 Flash (no thinking) · open | 1710 ±52 | 83.5 | 84.3 | 86.3 | 84.3 | 85.8 | 80.6 | 83.0 | 87.3 | 78.8 | 77.6 | $0.0033 | 0.2 min |
| #99 | DeepSeek V4 Flash (native chat alias) · open | 1495 ±60 | 80.6 | 80.6 | 83.4 | 81.6 | 82.5 | 74.0 | 78.8 | 84.1 | 83.5 | 78.2 | $0.0024 | 0.6 min |
| #108 | DeepSeek V4 Pro 0813 (default) · open | 1408 ±61 | 79.2 | 78.7 | 81.7 | 80.3 | 83.5 | 77.9 | 77.0 | 84.5 | 72.1 | 78.2 | $0.020 | 1.1 min |
| #111 | DeepSeek V4 Pro (xhigh) · open | 1370 ±56 | 79.0 | 80.0 | 81.4 | 79.5 | 82.5 | 75.2 | 77.5 | 80.2 | 78.5 | 70.8 | $0.044 | 1.8 min |
| #113 | DeepSeek V4 Pro (default) · open | 1326 ±39 | 78.7 | 79.3 | 81.5 | 79.8 | 83.2 | 75.6 | 78.7 | 78.7 | 74.3 | 70.5 | $0.030 | 1.0 min |
| #127 | DeepSeek V4 Flash (xhigh) · open | 1199 ±46 | 77.0 | 77.3 | 79.3 | 76.8 | 81.3 | 74.9 | 75.1 | 82.2 | 70.8 | 76.1 | $0.0033 | 3.0 min |
| #132 | DeepSeek V4 Flash (native reasoner alias) · open | 1173 ±45 | 75.9 | 75.4 | 79.9 | 78.2 | 80.8 | 69.4 | 75.1 | 81.4 | 66.3 | 78.8 | $0.0016 | 0.7 min |
| #134 | DeepSeek V3.2 · open | 1157 ±38 | 76.2 | 74.7 | 80.4 | 78.0 | 80.7 | 74.0 | 76.5 | 76.3 | 67.5 | 75.0 | $0.0040 | 1.1 min |
| #135 | DeepSeek V4 Flash (default) · open | 1150 ±40 | 75.8 | 75.1 | 79.7 | 77.7 | 81.3 | 69.5 | 75.2 | 81.8 | 65.6 | 79.1 | $0.0014 | 0.7 min |
I'd use DeepSeek V4.1 Flash (max). It's #27 on the whole board at $0.013 a script, which is cheap enough that most people can stop there. If you're generating at real volume, DeepSeek V4.1 Flash (default) is the value pick: 97 Elo behind, for a cheaper and faster script at $0.0077.
Here's the odd part. Pro is the bigger, pricier tier, and it doesn't pay off for scripts: its best run, DeepSeek V4 Pro 0813 (max), costs 3.5 times as much as DeepSeek V4.1 Flash (max) and only reaches #68. And don't bother with the bottom of the price list. DeepSeek V4 Flash (default) costs $0.0014, but it ranks #135, and when every option is this cheap, saving there makes no sense.
Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.
Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.
← All comparisons