Towards AITowards AIToneBench

Every OpenAI model, side by side

OpenAI is a hard family to pick from, just because there's so much of it: 49 configurations across 19 models, from GPT-4o mini up to GPT-6.1 Sol, most of the recent ones at several effort settings. Keeping Sol, Terra, Luna and Astra straight is a job in itself. You won't get the top script on the board here, but the top of the family is packed tight and cheap. The bottom is mostly older and smaller models, which is how you get 2015 Elo between first and last, so the OpenAI name alone tells you very little. The exact model and setting are what you're really picking.

Best writer
GPT-5.6 Sol (ultra)#31 · Elo 2148 · $0.249 per script
Best value (within 200 Elo of the best)
GPT-6 Sol (default)#55 · Elo 1955 · $0.054 per script
Cheapest
GPT-4o mini#174 · Elo 133 · $0.0025 per script

All 49 configurations

Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.

Board rankConfigurationEloOverallToneCraftSubstanceHookYouTubeFlowSlopLengthCuesCost / scriptTime / script
#31GPT-5.6 Sol (ultra)2148 ±3587.685.887.088.885.887.185.793.491.989.5$0.2493.3 min
#34GPT-6 Sol (max)2114 ±3887.284.786.288.586.287.485.894.391.288.3$0.1122.9 min
#35GPT-5.6 Sol (xhigh)2104 ±3487.285.687.187.985.886.585.692.292.188.9$0.1301.5 min
#36GPT-5.6 Sol (default)2100 ±3787.286.887.888.186.586.885.892.785.389.0$0.1061.1 min
#37GPT-5.6 Sol (none)2095 ±4087.286.787.787.686.486.685.892.487.288.8$0.1031.2 min
#38GPT-5.6 Sol (high)2093 ±3687.286.187.487.885.686.685.792.789.388.7$0.1111.4 min
#40GPT-5.6 Sol (low)2091 ±3787.186.587.787.986.586.685.892.285.788.8$0.1061.3 min
#47GPT-6 Sol (high)2037 ±3786.684.586.188.084.786.385.294.589.986.9$0.0601.2 min
#51GPT-5.6 Terra (xhigh)1976 ±3886.186.186.985.985.084.584.891.287.486.5$0.0661.5 min
#52GPT-5.6 Terra (default)1962 ±3986.086.887.385.585.884.284.990.584.486.0$0.0581.0 min
#53GPT-5.6 Terra (ultra)1959 ±4085.984.685.887.084.984.684.292.588.887.1$0.1593.8 min
#55GPT-6 Sol (default)1955 ±4385.884.186.087.884.486.185.194.482.886.7$0.0541.0 min
#56GPT-6.1 Sol (xhigh)1946 ±4485.781.984.287.383.985.483.194.694.787.1$0.1345.8 min
#57GPT-6.1 Sol (max)1944 ±5285.782.184.187.183.785.483.094.595.786.2$0.1788.0 min
#58GPT-5.6 Terra (none)1943 ±4285.886.587.285.285.584.484.690.884.185.9$0.0551.6 min
#59GPT-5.6 Terra (high)1939 ±4485.786.587.085.585.684.484.790.881.986.4$0.0591.0 min
#61GPT-6.1 Sol (low)1926 ±4285.682.385.386.985.186.284.394.686.787.4$0.0491.5 min
#62GPT-6 Astra (max)1921 ±4285.782.284.387.183.485.083.394.895.286.5$0.8528.8 min
#64GPT-5.6 Terra (low)1912 ±4285.486.586.985.685.083.684.590.481.086.1$0.0591.0 min
#66GPT-6.1 Sol (medium)1908 ±4585.382.084.786.585.086.083.994.386.587.7$0.0501.4 min
#67GPT-5.4 (xhigh)1888 ±6685.086.386.685.686.685.083.289.176.885.4$0.1923.8 min
#69GPT-6.1 Sol (high)1883 ±3685.481.984.986.984.185.883.694.388.787.4$0.0612.0 min
#72GPT-6 Astra (default)1837 ±4084.882.084.586.483.985.383.694.783.786.4$0.2441.4 min
#74GPT-5.6 Luna (ultra)1793 ±4084.382.284.286.484.384.381.692.084.286.2$0.0235.5 min
#78GPT-6 Luna (max)1695 ±4183.379.682.685.881.783.280.492.588.785.9$0.00412.0 min
#84GPT-5.5 (high)1621 ±5282.787.587.786.385.384.883.590.140.987.8$0.1781.5 min
#85GPT-6 Luna (high)1617 ±4682.479.783.286.081.978.280.291.684.284.8$0.00281.1 min
#86GPT-5.5 (xhigh)1612 ±4182.687.587.686.285.185.083.489.741.087.2$0.1781.4 min
#87GPT-6 Luna (default)1601 ±4182.379.983.285.182.076.580.191.587.684.4$0.00251.1 min
#89GPT-5.6 Luna (xhigh)1600 ±4382.482.084.685.883.882.581.591.266.385.8$0.00921.9 min
#93GPT-5.41559 ±4382.085.585.483.484.382.280.687.160.284.2$0.0801.1 min
#94GPT-5.5 (default)1550 ±4181.987.687.686.385.585.382.989.732.187.3$0.1711.4 min
#95GPT-5.6 Luna (high)1534 ±4681.782.384.885.583.981.181.090.759.685.1$0.00711.3 min
#97GPT-5.4 mini (high)1523 ±4381.483.684.081.082.779.578.586.976.979.7$0.0532.4 min
#98GPT-5.6 Luna (none)1513 ±4481.482.484.884.683.581.580.589.659.285.0$0.00611.0 min
#101GPT-5.5 (none)1492 ±3981.287.787.386.285.985.282.888.924.187.5$0.1721.3 min
#102GPT-5.6 Luna (default)1477 ±3781.083.084.984.984.281.780.990.350.585.4$0.00651.1 min
#105GPT-5.6 Luna (low)1447 ±2980.782.284.585.083.280.580.489.452.685.0$0.00631.1 min
#107GPT-5.4 mini1447 ±4180.483.283.780.982.077.877.886.271.978.4$0.0361.5 min
#118GPT-5.21282 ±4078.686.584.785.085.881.779.086.316.184.7$0.0801.2 min
#120GPT-51259 ±4078.083.482.881.285.081.172.186.540.583.3$0.0731.0 min
#136GPT-5.11145 ±3276.182.281.583.084.478.877.286.412.885.0$0.0540.7 min
#137GPT-5 mini1138 ±4375.874.475.275.179.573.867.183.889.278.6$0.0110.9 min
#159GPT-5.4 nano751 ±3469.075.074.278.777.265.765.481.817.176.3$0.00710.3 min
#164gpt-oss 120B (high) · open648 ±3767.266.269.860.474.168.465.072.170.962.1$0.00631.0 min
#165gpt-oss 120B · open638 ±4167.266.670.362.774.666.366.171.567.959.1$0.00330.4 min
#170GPT-5 nano321 ±5559.359.358.762.968.660.044.968.457.365.4$0.00360.5 min
#173GPT-4o183 ±3056.657.161.259.365.253.557.356.933.959.6$0.0400.2 min
#174GPT-4o mini133 ±3255.754.861.163.064.246.555.459.445.538.4$0.00250.2 min

Which one should you use?

If you only want the best OpenAI script, that's GPT-5.6 Sol (ultra) at #31, for $0.249 per script. I'd still try GPT-6 Sol (max) first, though: it sits 34 Elo behind, the two ranges overlap, and the ultra setting costs 2.2 times as much.

For volume, GPT-6 Sol (default) is the value pick, 193 Elo off the top for $0.054 a script. On a tight budget, I'd go with GPT-6 Luna (max): $0.0041 gets you #78 on the board, which is a lot of writer for that price. The one I'd skip is GPT-6 Astra (max). It costs 7.6 times what GPT-6 Sol (max) does and ranks lower, so for writing, at least, I don't see where it fits.

Thinking levels, model by model

Each of these models ran at more than one reasoning-effort setting. Their pages show what each step up buys.

Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.

← All comparisons