The same model can land 39 places apart on this board just by changing its effort setting. That's why every card here is one configuration, not one model name. Each card gives you the rank, writing Elo, overall score and cost per script, and a click opens the full scorecard. 167 configurations are on the board, and every one of them has all 10 scripts scored and a current rank.
167 configurations from more than 13 labs are on the board right now, and every one of them has all 10 scripts scored and a current rank. A configuration is one model at one setting, so the same model at low and at max effort counts twice, and it should: effort alone can move a model a long way. You'll find Anthropic, OpenAI, Google, Meta, DeepSeek, xAI, Mistral, Alibaba (Qwen), Moonshot (Kimi), Zhipu (GLM) and more, closed and open weights alike. Every ranked configuration writes the scripts behind the same 10 of my What's AI videos, and each draft is scored blind against my finished version. Archived rows stay visible, but they never enter the current rank.
Claude Opus 5.5 (max effort) is #1 right now, with a writing Elo of 2846 and 91.8/100 overall. It also costs $3.43 per script, so if you're writing at volume, read the value answer below before you wire it into a pipeline. Its scorecard breaks the result down by metric and by script, and the leaderboard has the whole ranking.
GLM-5.3 is the strongest open-weights writer here: #15, Elo 2256, at $0.114 per script. That's 590 Elo behind the leader, so the closed frontier still writes better. But if you need to run the model yourself or keep your data in-house, this is the one to try first.
GLM-5.3 Flash is the strongest writer in the cheaper half of the board: Elo 2244 at $0.007 per script. When cost matters, that's where I'd start. Each card shows its cost, and you can sort the leaderboard by any metric.
53 of the 167 tracked configurations have open weights; the rest are proprietary. The filter buttons at the top split the grid so you can look at each group on its own.
New models go on the board as soon as practical after release, and new scripts get added a few times a year. New scripts are normally benchmarked before their video goes out, though that doesn't make it impossible for a model to have seen one. Every card and score regenerates from the live board, and the methodology explains how the scoring works.
Yes. Click any card for that configuration's scorecard: stat cards, a written recap, all 9 metric scores with their board rank, per-script results, and its own FAQ.