The leaderboard tells you who's first. It doesn't tell you whether first place is worth the price gap over the model you already use, and for most of us that's the real question. So among the labs this page tracks by name, I took the 8 whose best configuration ranks highest on Elo (Anthropic, Zhipu, Moonshot, OpenAI, xAI, DeepSeek, Alibaba and MiniMax), and put their best configurations side by side on the same 10 real scripts, under the same three-judge panel: overall score, Elo, all nine metrics, cost per script, consistency, and a verdict on which one to pick. That's 28 match-ups. The highest-ranked model outside the grid right now is MiMo-V2.6-Pro UltraSpeed (#24): its lab isn't on the list this page tracks yet. Every ranked model still gets a full scorecard on the models page.
Pairs from earlier boards. Some of these models have since lost their spot as their lab's best, but the numbers refresh with every rebuild, and a model without complete current coverage shows N/A instead of an old score.