OpenAI is a hard family to pick from, just because there's so much of it: 49 configurations across 19 models, from GPT-4o mini up to GPT-6.1 Sol, most of the recent ones at several effort settings. Keeping Sol, Terra, Luna and Astra straight is a job in itself. You won't get the top script on the board here, but the top of the family is packed tight and cheap. The bottom is mostly older and smaller models, which is how you get 2015 Elo between first and last, so the OpenAI name alone tells you very little. The exact model and setting are what you're really picking.
Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.
| Board rank | Configuration | Elo | Overall | Tone | Craft | Substance | Hook | YouTube | Flow | Slop | Length | Cues | Cost / script | Time / script |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| #31 | GPT-5.6 Sol (ultra) | 2148 ±35 | 87.6 | 85.8 | 87.0 | 88.8 | 85.8 | 87.1 | 85.7 | 93.4 | 91.9 | 89.5 | $0.249 | 3.3 min |
| #34 | GPT-6 Sol (max) | 2114 ±38 | 87.2 | 84.7 | 86.2 | 88.5 | 86.2 | 87.4 | 85.8 | 94.3 | 91.2 | 88.3 | $0.112 | 2.9 min |
| #35 | GPT-5.6 Sol (xhigh) | 2104 ±34 | 87.2 | 85.6 | 87.1 | 87.9 | 85.8 | 86.5 | 85.6 | 92.2 | 92.1 | 88.9 | $0.130 | 1.5 min |
| #36 | GPT-5.6 Sol (default) | 2100 ±37 | 87.2 | 86.8 | 87.8 | 88.1 | 86.5 | 86.8 | 85.8 | 92.7 | 85.3 | 89.0 | $0.106 | 1.1 min |
| #37 | GPT-5.6 Sol (none) | 2095 ±40 | 87.2 | 86.7 | 87.7 | 87.6 | 86.4 | 86.6 | 85.8 | 92.4 | 87.2 | 88.8 | $0.103 | 1.2 min |
| #38 | GPT-5.6 Sol (high) | 2093 ±36 | 87.2 | 86.1 | 87.4 | 87.8 | 85.6 | 86.6 | 85.7 | 92.7 | 89.3 | 88.7 | $0.111 | 1.4 min |
| #40 | GPT-5.6 Sol (low) | 2091 ±37 | 87.1 | 86.5 | 87.7 | 87.9 | 86.5 | 86.6 | 85.8 | 92.2 | 85.7 | 88.8 | $0.106 | 1.3 min |
| #47 | GPT-6 Sol (high) | 2037 ±37 | 86.6 | 84.5 | 86.1 | 88.0 | 84.7 | 86.3 | 85.2 | 94.5 | 89.9 | 86.9 | $0.060 | 1.2 min |
| #51 | GPT-5.6 Terra (xhigh) | 1976 ±38 | 86.1 | 86.1 | 86.9 | 85.9 | 85.0 | 84.5 | 84.8 | 91.2 | 87.4 | 86.5 | $0.066 | 1.5 min |
| #52 | GPT-5.6 Terra (default) | 1962 ±39 | 86.0 | 86.8 | 87.3 | 85.5 | 85.8 | 84.2 | 84.9 | 90.5 | 84.4 | 86.0 | $0.058 | 1.0 min |
| #53 | GPT-5.6 Terra (ultra) | 1959 ±40 | 85.9 | 84.6 | 85.8 | 87.0 | 84.9 | 84.6 | 84.2 | 92.5 | 88.8 | 87.1 | $0.159 | 3.8 min |
| #55 | GPT-6 Sol (default) | 1955 ±43 | 85.8 | 84.1 | 86.0 | 87.8 | 84.4 | 86.1 | 85.1 | 94.4 | 82.8 | 86.7 | $0.054 | 1.0 min |
| #56 | GPT-6.1 Sol (xhigh) | 1946 ±44 | 85.7 | 81.9 | 84.2 | 87.3 | 83.9 | 85.4 | 83.1 | 94.6 | 94.7 | 87.1 | $0.134 | 5.8 min |
| #57 | GPT-6.1 Sol (max) | 1944 ±52 | 85.7 | 82.1 | 84.1 | 87.1 | 83.7 | 85.4 | 83.0 | 94.5 | 95.7 | 86.2 | $0.178 | 8.0 min |
| #58 | GPT-5.6 Terra (none) | 1943 ±42 | 85.8 | 86.5 | 87.2 | 85.2 | 85.5 | 84.4 | 84.6 | 90.8 | 84.1 | 85.9 | $0.055 | 1.6 min |
| #59 | GPT-5.6 Terra (high) | 1939 ±44 | 85.7 | 86.5 | 87.0 | 85.5 | 85.6 | 84.4 | 84.7 | 90.8 | 81.9 | 86.4 | $0.059 | 1.0 min |
| #61 | GPT-6.1 Sol (low) | 1926 ±42 | 85.6 | 82.3 | 85.3 | 86.9 | 85.1 | 86.2 | 84.3 | 94.6 | 86.7 | 87.4 | $0.049 | 1.5 min |
| #62 | GPT-6 Astra (max) | 1921 ±42 | 85.7 | 82.2 | 84.3 | 87.1 | 83.4 | 85.0 | 83.3 | 94.8 | 95.2 | 86.5 | $0.852 | 8.8 min |
| #64 | GPT-5.6 Terra (low) | 1912 ±42 | 85.4 | 86.5 | 86.9 | 85.6 | 85.0 | 83.6 | 84.5 | 90.4 | 81.0 | 86.1 | $0.059 | 1.0 min |
| #66 | GPT-6.1 Sol (medium) | 1908 ±45 | 85.3 | 82.0 | 84.7 | 86.5 | 85.0 | 86.0 | 83.9 | 94.3 | 86.5 | 87.7 | $0.050 | 1.4 min |
| #67 | GPT-5.4 (xhigh) | 1888 ±66 | 85.0 | 86.3 | 86.6 | 85.6 | 86.6 | 85.0 | 83.2 | 89.1 | 76.8 | 85.4 | $0.192 | 3.8 min |
| #69 | GPT-6.1 Sol (high) | 1883 ±36 | 85.4 | 81.9 | 84.9 | 86.9 | 84.1 | 85.8 | 83.6 | 94.3 | 88.7 | 87.4 | $0.061 | 2.0 min |
| #72 | GPT-6 Astra (default) | 1837 ±40 | 84.8 | 82.0 | 84.5 | 86.4 | 83.9 | 85.3 | 83.6 | 94.7 | 83.7 | 86.4 | $0.244 | 1.4 min |
| #74 | GPT-5.6 Luna (ultra) | 1793 ±40 | 84.3 | 82.2 | 84.2 | 86.4 | 84.3 | 84.3 | 81.6 | 92.0 | 84.2 | 86.2 | $0.023 | 5.5 min |
| #78 | GPT-6 Luna (max) | 1695 ±41 | 83.3 | 79.6 | 82.6 | 85.8 | 81.7 | 83.2 | 80.4 | 92.5 | 88.7 | 85.9 | $0.0041 | 2.0 min |
| #84 | GPT-5.5 (high) | 1621 ±52 | 82.7 | 87.5 | 87.7 | 86.3 | 85.3 | 84.8 | 83.5 | 90.1 | 40.9 | 87.8 | $0.178 | 1.5 min |
| #85 | GPT-6 Luna (high) | 1617 ±46 | 82.4 | 79.7 | 83.2 | 86.0 | 81.9 | 78.2 | 80.2 | 91.6 | 84.2 | 84.8 | $0.0028 | 1.1 min |
| #86 | GPT-5.5 (xhigh) | 1612 ±41 | 82.6 | 87.5 | 87.6 | 86.2 | 85.1 | 85.0 | 83.4 | 89.7 | 41.0 | 87.2 | $0.178 | 1.4 min |
| #87 | GPT-6 Luna (default) | 1601 ±41 | 82.3 | 79.9 | 83.2 | 85.1 | 82.0 | 76.5 | 80.1 | 91.5 | 87.6 | 84.4 | $0.0025 | 1.1 min |
| #89 | GPT-5.6 Luna (xhigh) | 1600 ±43 | 82.4 | 82.0 | 84.6 | 85.8 | 83.8 | 82.5 | 81.5 | 91.2 | 66.3 | 85.8 | $0.0092 | 1.9 min |
| #93 | GPT-5.4 | 1559 ±43 | 82.0 | 85.5 | 85.4 | 83.4 | 84.3 | 82.2 | 80.6 | 87.1 | 60.2 | 84.2 | $0.080 | 1.1 min |
| #94 | GPT-5.5 (default) | 1550 ±41 | 81.9 | 87.6 | 87.6 | 86.3 | 85.5 | 85.3 | 82.9 | 89.7 | 32.1 | 87.3 | $0.171 | 1.4 min |
| #95 | GPT-5.6 Luna (high) | 1534 ±46 | 81.7 | 82.3 | 84.8 | 85.5 | 83.9 | 81.1 | 81.0 | 90.7 | 59.6 | 85.1 | $0.0071 | 1.3 min |
| #97 | GPT-5.4 mini (high) | 1523 ±43 | 81.4 | 83.6 | 84.0 | 81.0 | 82.7 | 79.5 | 78.5 | 86.9 | 76.9 | 79.7 | $0.053 | 2.4 min |
| #98 | GPT-5.6 Luna (none) | 1513 ±44 | 81.4 | 82.4 | 84.8 | 84.6 | 83.5 | 81.5 | 80.5 | 89.6 | 59.2 | 85.0 | $0.0061 | 1.0 min |
| #101 | GPT-5.5 (none) | 1492 ±39 | 81.2 | 87.7 | 87.3 | 86.2 | 85.9 | 85.2 | 82.8 | 88.9 | 24.1 | 87.5 | $0.172 | 1.3 min |
| #102 | GPT-5.6 Luna (default) | 1477 ±37 | 81.0 | 83.0 | 84.9 | 84.9 | 84.2 | 81.7 | 80.9 | 90.3 | 50.5 | 85.4 | $0.0065 | 1.1 min |
| #105 | GPT-5.6 Luna (low) | 1447 ±29 | 80.7 | 82.2 | 84.5 | 85.0 | 83.2 | 80.5 | 80.4 | 89.4 | 52.6 | 85.0 | $0.0063 | 1.1 min |
| #107 | GPT-5.4 mini | 1447 ±41 | 80.4 | 83.2 | 83.7 | 80.9 | 82.0 | 77.8 | 77.8 | 86.2 | 71.9 | 78.4 | $0.036 | 1.5 min |
| #118 | GPT-5.2 | 1282 ±40 | 78.6 | 86.5 | 84.7 | 85.0 | 85.8 | 81.7 | 79.0 | 86.3 | 16.1 | 84.7 | $0.080 | 1.2 min |
| #120 | GPT-5 | 1259 ±40 | 78.0 | 83.4 | 82.8 | 81.2 | 85.0 | 81.1 | 72.1 | 86.5 | 40.5 | 83.3 | $0.073 | 1.0 min |
| #136 | GPT-5.1 | 1145 ±32 | 76.1 | 82.2 | 81.5 | 83.0 | 84.4 | 78.8 | 77.2 | 86.4 | 12.8 | 85.0 | $0.054 | 0.7 min |
| #137 | GPT-5 mini | 1138 ±43 | 75.8 | 74.4 | 75.2 | 75.1 | 79.5 | 73.8 | 67.1 | 83.8 | 89.2 | 78.6 | $0.011 | 0.9 min |
| #159 | GPT-5.4 nano | 751 ±34 | 69.0 | 75.0 | 74.2 | 78.7 | 77.2 | 65.7 | 65.4 | 81.8 | 17.1 | 76.3 | $0.0071 | 0.3 min |
| #164 | gpt-oss 120B (high) · open | 648 ±37 | 67.2 | 66.2 | 69.8 | 60.4 | 74.1 | 68.4 | 65.0 | 72.1 | 70.9 | 62.1 | $0.0063 | 1.0 min |
| #165 | gpt-oss 120B · open | 638 ±41 | 67.2 | 66.6 | 70.3 | 62.7 | 74.6 | 66.3 | 66.1 | 71.5 | 67.9 | 59.1 | $0.0033 | 0.4 min |
| #170 | GPT-5 nano | 321 ±55 | 59.3 | 59.3 | 58.7 | 62.9 | 68.6 | 60.0 | 44.9 | 68.4 | 57.3 | 65.4 | $0.0036 | 0.5 min |
| #173 | GPT-4o | 183 ±30 | 56.6 | 57.1 | 61.2 | 59.3 | 65.2 | 53.5 | 57.3 | 56.9 | 33.9 | 59.6 | $0.040 | 0.2 min |
| #174 | GPT-4o mini | 133 ±32 | 55.7 | 54.8 | 61.1 | 63.0 | 64.2 | 46.5 | 55.4 | 59.4 | 45.5 | 38.4 | $0.0025 | 0.2 min |
If you only want the best OpenAI script, that's GPT-5.6 Sol (ultra) at #31, for $0.249 per script. I'd still try GPT-6 Sol (max) first, though: it sits 34 Elo behind, the two ranges overlap, and the ultra setting costs 2.2 times as much.
For volume, GPT-6 Sol (default) is the value pick, 193 Elo off the top for $0.054 a script. On a tight budget, I'd go with GPT-6 Luna (max): $0.0041 gets you #78 on the board, which is a lot of writer for that price. The one I'd skip is GPT-6 Astra (max). It costs 7.6 times what GPT-6 Sol (max) does and ranks lower, so for writing, at least, I don't see where it fits.
Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.
Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.
← All comparisons