130 configurations tracked; 130 with complete 9-task coverage; 130 with current ranks.
130 model configurations are tracked across 14 providers; 130 have complete coverage for all 9 current tasks, while 130 are eligible for and form the current rank. The catalog spans Anthropic, OpenAI, Google, Meta, DeepSeek, xAI, Mistral, Alibaba/Qwen, Moonshot/Kimi, Zhipu/GLM and more, across proprietary and open weights. Every current-ranked model writes the same 9 real YouTube scripts and is scored blind against our finished versions; archived rows are not compared in the current rank.
Claude Opus 5 (max effort) tops the board with a writing Elo of 2649 and 90.8/100 overall. Open its scorecard for the full per-metric and per-article breakdown, or see the leaderboard for the ranking.
Kimi K3 is the strongest open-weights model here, at Elo 2438. Open weights still trail the closed frontier on pure voice fidelity, but they have closed a lot of the gap at a fraction of the cost.
DeepSeek V4 Flash 0731 gives the most writing quality per dollar in the cheaper half of the board (Elo 1955 at $0.005/task). Use the cost column on each card, or sort the leaderboard by any metric.
47 of the 130 tracked configurations are open-weights; the rest are proprietary. Use the filter buttons above to see each group on its own.
New models are added as soon as practical after release (we monitor the major provider APIs daily), and new articles are added a few times a year. Source-publication timing is recorded, including disclosed exceptions to the normal pre-publication workflow. Every card and score regenerates from the live board. See the methodology for how scoring works.
Yes. Click any card to open that model's scorecard: stat cards, a written recap, all nine metric scores with board rank, per-article results, and its own FAQ.