Towards AITowards AIToneBench

Thinking levels · OpenAI

Does GPT-6 Luna write better when it thinks harder?

A bit, and at Luna's prices the extra thinking is close to free. Going from high to max adds 78 Elo, enough to win about 61% of script comparisons. What you actually pay is time: 2.0 minutes a script instead of 1.1.

X axis

Every setting, step by step

SettingEloStepOverallCost / scriptTime / scriptBoard rank
high1617 ±46-82.4$0.00281.1 min#85
max1695 ±41+7883.3$0.00412.0 min#78
default (no effort flag)1601 ±41-82.3$0.00251.1 min#87

The 95% Elo ranges overlap: GPT-6 Luna (high) runs from 1570 to 1662, and GPT-6 Luna (max) from 1654 to 1736.

Which setting should you use?

Here I'd just run max. It's the only setting that pulls away from the default you get with no flag, mostly on the YouTube best-practices score. When the whole difference is a fraction of a cent, there's no reason to save it, right? Just keep your expectations in check: GPT-6 Luna (max) is at #78, so treat it as a cheap drafting model and plan a real editing pass before anything goes out.

Other OpenAI ladders

Elo comes from comparing every pair of models on every script across the full board, so these settings sit on the same scale as the leaderboard. "Step" is the Elo change from the previous explicit setting. Cost is one full script at API list prices, and time is one accepted script, retries included. Compare this ladder with other models on the thinking-levels page, or see every OpenAI model side by side.