Towards AITowards AIToneBench

Thinking level analysis

Does thinking harder write better?

Writing score against cost per script for the 38 models we benchmarked at more than one reasoning-effort setting (107 settings, 10 scripts × 5 runs each, scored blind by three judges). Each line walks one model from its lowest to its highest effort; a hollow marker is the default you get when you pass no effort flag.

X axis

Elo comes from all pairwise draft comparisons on the full board; cost is the measured generation cost per script at list prices; time is the mean wall-clock per script. Select a point for the model's scorecard, or read the methodology.