Towards AITowards AIToneBench

Head-to-head comparisons

The leaderboard tells you who's first. It doesn't tell you whether first place is worth the price gap over the model you already use, and for most of us that's the real question. So among the labs this page tracks by name, I took the 8 whose best configuration ranks highest on Elo (Anthropic, Zhipu, Moonshot, OpenAI, xAI, DeepSeek, Alibaba and MiniMax), and put their best configurations side by side on the same 10 real scripts, under the same three-judge panel: overall score, Elo, all nine metrics, cost per script, consistency, and a verdict on which one to pick. That's 28 match-ups. The highest-ranked model outside the grid right now is MiMo-V2.6-Pro UltraSpeed (#24): its lab isn't on the list this page tracks yet. Every ranked model still gets a full scorecard on the models page.

Claude Opus 5.5 (max) vs GLM-5.3
91.8 vs 88.7
Claude Opus 5.5 (max) vs Kimi K3
91.8 vs 88.5
Claude Opus 5.5 (max) vs GPT-5.6 Sol (ultra)
91.8 vs 88.0
Claude Opus 5.5 (max) vs Grok 4.7 (high)
91.8 vs 87.8
Claude Opus 5.5 (max) vs DeepSeek V4.1 Flash (max)
91.8 vs 88.1
Claude Opus 5.5 (max) vs Qwen3.8 Flash
91.8 vs 83.0
Claude Opus 5.5 (max) vs MiniMax M3
91.8 vs 82.7
GLM-5.3 vs Kimi K3
88.7 vs 88.5
GLM-5.3 vs GPT-5.6 Sol (ultra)
88.7 vs 88.0
GLM-5.3 vs Grok 4.7 (high)
88.7 vs 87.8
GLM-5.3 vs DeepSeek V4.1 Flash (max)
88.7 vs 88.1
GLM-5.3 vs Qwen3.8 Flash
88.7 vs 83.0
GLM-5.3 vs MiniMax M3
88.7 vs 82.7
Kimi K3 vs GPT-5.6 Sol (ultra)
88.5 vs 88.0
Kimi K3 vs Grok 4.7 (high)
88.5 vs 87.8
Kimi K3 vs DeepSeek V4.1 Flash (max)
88.5 vs 88.1
Kimi K3 vs Qwen3.8 Flash
88.5 vs 83.0
Kimi K3 vs MiniMax M3
88.5 vs 82.7
GPT-5.6 Sol (ultra) vs Grok 4.7 (high)
88.0 vs 87.8
GPT-5.6 Sol (ultra) vs DeepSeek V4.1 Flash (max)
88.0 vs 88.1
GPT-5.6 Sol (ultra) vs Qwen3.8 Flash
88.0 vs 83.0
GPT-5.6 Sol (ultra) vs MiniMax M3
88.0 vs 82.7
Grok 4.7 (high) vs DeepSeek V4.1 Flash (max)
87.8 vs 88.1
Grok 4.7 (high) vs Qwen3.8 Flash
87.8 vs 83.0
Grok 4.7 (high) vs MiniMax M3
87.8 vs 82.7
DeepSeek V4.1 Flash (max) vs Qwen3.8 Flash
88.1 vs 83.0
DeepSeek V4.1 Flash (max) vs MiniMax M3
88.1 vs 82.7
Qwen3.8 Flash vs MiniMax M3
83.0 vs 82.7

Past match-ups

Pairs from earlier boards. Some of these models have since lost their spot as their lab's best, but the numbers refresh with every rebuild, and a model without complete current coverage shows N/A instead of an old score.

Claude Fable 5.1 (max) vs Kimi K3
90.5 vs 88.5
Claude Fable 5.1 (max) vs DeepSeek V4.1 Flash (max)
90.5 vs 88.1
Claude Fable 5.1 (max) vs GLM-5.3
90.5 vs 88.7
Claude Fable 5.1 (max) vs Grok 4.6
90.5 vs 86.3
Claude Fable 5.1 (max) vs MiniMax M3
90.5 vs 82.7
Claude Fable 5.1 (max) vs Qwen3.8 Max
90.5 vs 83.2
Kimi K3 vs Grok 4.6
88.5 vs 86.3
Kimi K3 vs Qwen3.8 Max
88.5 vs 83.2
GPT-5.6 Sol (ultra) vs Qwen3.8 Max
88.0 vs 83.2
DeepSeek V4.1 Flash (max) vs Grok 4.6
88.1 vs 86.3
DeepSeek V4.1 Flash (max) vs Qwen3.8 Max
88.1 vs 83.2
GLM-5.3 vs Grok 4.6
88.7 vs 86.3
GLM-5.3 vs Qwen3.8 Max
88.7 vs 83.2
Grok 4.6 vs MiniMax M3
86.3 vs 82.7
Grok 4.6 vs Qwen3.8 Max
86.3 vs 83.2
Claude Fable 5 (max) vs Kimi K3 (thinking)
89.7 vs 87.2
Claude Fable 5 (max) vs GPT-5.6 Sol (high)
89.7 vs 87.6
Claude Fable 5 (max) vs Grok 4.5
89.7 vs 81.9
Claude Fable 5 (max) vs GLM-5
89.7 vs 82.8
Claude Fable 5 (max) vs MiniMax M3
89.7 vs 82.7
Claude Fable 5 (max) vs Qwen3.7 Max (high)
89.7 vs 80.8
Claude Fable 5 (max) vs Gemini 3.1 Pro (default)
89.7 vs 80.1
Kimi K3 (thinking) vs Grok 4.5
87.2 vs 81.9
Kimi K3 (thinking) vs GLM-5
87.2 vs 82.8
Kimi K3 (thinking) vs MiniMax M3
87.2 vs 82.7
Kimi K3 (thinking) vs Qwen3.7 Max (high)
87.2 vs 80.8
Kimi K3 (thinking) vs Gemini 3.1 Pro (default)
87.2 vs 80.1
GPT-5.6 Sol (high) vs Grok 4.5
87.6 vs 81.9
GPT-5.6 Sol (high) vs GLM-5
87.6 vs 82.8
GPT-5.6 Sol (high) vs MiniMax M3
87.6 vs 82.7
GPT-5.6 Sol (high) vs Qwen3.7 Max (high)
87.6 vs 80.8
GPT-5.6 Sol (high) vs Gemini 3.1 Pro (default)
87.6 vs 80.1
Grok 4.5 vs GLM-5
81.9 vs 82.8
Grok 4.5 vs MiniMax M3
81.9 vs 82.7
Grok 4.5 vs Qwen3.7 Max (high)
81.9 vs 80.8
Grok 4.5 vs Gemini 3.1 Pro (default)
81.9 vs 80.1
GLM-5 vs Qwen3.7 Max (high)
82.8 vs 80.8
GLM-5 vs Gemini 3.1 Pro (default)
82.8 vs 80.1
MiniMax M3 vs Qwen3.7 Max (high)
82.7 vs 80.8
MiniMax M3 vs Gemini 3.1 Pro (default)
82.7 vs 80.1
Claude Fable 5 (max) vs DeepSeek V4 Pro (xhigh)
89.7 vs 78.9
Kimi K3 (thinking) vs DeepSeek V4 Pro (xhigh)
87.2 vs 78.9
GPT-5.6 Sol (high) vs DeepSeek V4 Pro (xhigh)
87.6 vs 78.9
Grok 4.5 vs DeepSeek V4 Pro (xhigh)
81.9 vs 78.9
GLM-5 vs DeepSeek V4 Pro (xhigh)
82.8 vs 78.9
MiniMax M3 vs DeepSeek V4 Pro (xhigh)
82.7 vs 78.9
Qwen3.7 Max (high) vs DeepSeek V4 Pro (xhigh)
80.8 vs 78.9
Claude Opus 5 (max) vs Kimi K3
89.4 vs 88.5
Claude Opus 5 (max) vs Kimi K3 (thinking)
89.4 vs 87.2
Claude Opus 5 (max) vs GPT-5.6 Sol (high)
89.4 vs 87.6
Claude Opus 5 (max) vs Grok 4.5
89.4 vs 81.9
Claude Opus 5 (max) vs Gemini 3.1 Pro (default)
89.4 vs 80.1
Claude Opus 5 (max) vs GLM-5
89.4 vs 82.8
Claude Opus 5 (max) vs MiniMax M3
89.4 vs 82.7
Claude Opus 5 (max) vs Qwen3.7 Max (default)
89.4 vs 80.9
Claude Opus 5 (max) vs Qwen3.7 Max (high)
89.4 vs 80.8
Claude Opus 5 (max) vs DeepSeek V4 Pro (xhigh)
89.4 vs 78.9
Kimi K3 vs GPT-5.6 Sol (high)
88.5 vs 87.6
Kimi K3 vs Grok 4.5
88.5 vs 81.9
Kimi K3 vs GLM-5
88.5 vs 82.8
Kimi K3 vs Qwen3.7 Max (high)
88.5 vs 80.8
Kimi K3 vs Gemini 3.1 Pro (default)
88.5 vs 80.1
Kimi K3 (thinking) vs Qwen3.7 Max (default)
87.2 vs 80.9
GPT-5.6 Sol (high) vs Qwen3.7 Max (default)
87.6 vs 80.9
Grok 4.5 vs Qwen3.7 Max (default)
81.9 vs 80.9
GLM-5 vs Qwen3.7 Max (default)
82.8 vs 80.9
MiniMax M3 vs Qwen3.7 Max (default)
82.7 vs 80.9
Qwen3.7 Max (default) vs DeepSeek V4 Pro (xhigh)
80.9 vs 78.9
Claude Opus 5 (max) vs GPT-5.6 Sol (ultra)
89.4 vs 88.0
Claude Opus 5 (max) vs Qwen3.8 Max
89.4 vs 83.2
Claude Opus 5 (max) vs DeepSeek V4 Flash 0731
89.4 vs 84.4
GPT-5.6 Sol (ultra) vs Kimi K3 (thinking)
88.0 vs 87.2
GPT-5.6 Sol (ultra) vs Grok 4.5
88.0 vs 81.9
GPT-5.6 Sol (ultra) vs DeepSeek V4 Flash 0731
88.0 vs 84.4
GPT-5.6 Sol (ultra) vs GLM-5
88.0 vs 82.8
Kimi K3 (thinking) vs Qwen3.8 Max
87.2 vs 83.2
Kimi K3 (thinking) vs DeepSeek V4 Flash 0731
87.2 vs 84.4
Grok 4.5 vs Qwen3.8 Max
81.9 vs 83.2
Grok 4.5 vs DeepSeek V4 Flash 0731
81.9 vs 84.4
Qwen3.8 Max vs DeepSeek V4 Flash 0731
83.2 vs 84.4
DeepSeek V4 Flash 0731 vs GLM-5
84.4 vs 82.8
DeepSeek V4 Flash 0731 vs MiniMax M3
84.4 vs 82.7
Kimi K3 vs DeepSeek V4 Flash 0731
88.5 vs 84.4
Claude Opus 5 (max) vs Grok 4.6 (high)
89.4 vs 85.9
GPT-5.6 Sol (ultra) vs Grok 4.6 (high)
88.0 vs 85.9
Grok 4.6 (high) vs DeepSeek V4 Flash 0731
85.9 vs 84.4
Grok 4.6 (high) vs GLM-5
85.9 vs 82.8
Grok 4.6 (high) vs MiniMax M3
85.9 vs 82.7
Grok 4.6 (high) vs Qwen3.8 Max
85.9 vs 83.2
Kimi K3 vs Grok 4.6 (high)
88.5 vs 85.9
Claude Opus 5 (max) vs DeepSeek V4 Pro 0813 (max)
89.4 vs 84.8
DeepSeek V4 Pro 0813 (max) vs GLM-5
84.8 vs 82.8
DeepSeek V4 Pro 0813 (max) vs MiniMax M3
84.8 vs 82.7
DeepSeek V4 Pro 0813 (max) vs Qwen3.8 Max
84.8 vs 83.2
GPT-5.6 Sol (high) vs DeepSeek V4 Pro 0813 (max)
87.6 vs 84.8
GPT-5.6 Sol (high) vs Qwen3.8 Max
87.6 vs 83.2
Grok 4.6 (high) vs DeepSeek V4 Pro 0813 (max)
85.9 vs 84.8
Kimi K3 vs DeepSeek V4 Pro 0813 (max)
88.5 vs 84.8
Claude Fable 5 (max) vs DeepSeek V4 Pro 0813 (max)
89.7 vs 84.8
Claude Fable 5 (max) vs GPT-5.6 Sol (ultra)
89.7 vs 88.0
Claude Fable 5 (max) vs Grok 4.6
89.7 vs 86.3
Claude Fable 5 (max) vs Kimi K3
89.7 vs 88.5
Claude Fable 5 (max) vs Qwen3.8 Max
89.7 vs 83.2
GPT-5.6 Sol (ultra) vs DeepSeek V4 Pro 0813 (max)
88.0 vs 84.8
Grok 4.6 vs DeepSeek V4 Pro 0813 (max)
86.3 vs 84.8
Grok 4.6 vs GLM-5
86.3 vs 82.8
Claude Fable 5 (max) vs GLM-5.3
89.7 vs 88.7
GLM-5.3 vs DeepSeek V4 Pro 0813 (max)
88.7 vs 84.8
Claude Fable 5 (max) vs Muse Spark 1.3 (thinking)
89.7 vs 82.8
DeepSeek V4 Pro 0813 (max) vs Muse Spark 1.3 (thinking)
84.8 vs 82.8
GLM-5.3 vs Muse Spark 1.3 (thinking)
88.7 vs 82.8
GPT-5.6 Sol (ultra) vs Muse Spark 1.3 (thinking)
88.0 vs 82.8
Grok 4.6 vs Muse Spark 1.3 (thinking)
86.3 vs 82.8
Kimi K3 vs Muse Spark 1.3 (thinking)
88.5 vs 82.8
MiniMax M3 vs Muse Spark 1.3 (thinking)
82.7 vs 82.8
Claude Opus 5.5 (max) vs GPT-6 Sol (max)
91.8 vs 88.0
Claude Opus 5.5 (max) vs Qwen3.8 Max
91.8 vs 83.2
GLM-5.3 vs GPT-6 Sol (max)
88.7 vs 88.0
GPT-6 Sol (max) vs MiniMax M3
88.0 vs 82.7
GPT-6 Sol (max) vs Qwen3.8 Max
88.0 vs 83.2
Grok 4.7 (high) vs Qwen3.8 Max
87.8 vs 83.2
GLM-5 vs MiniMax M3
82.8 vs 82.7
GLM-5 vs Qwen3.8 Max
82.8 vs 83.2
DeepSeek V4.1 Flash (max) vs GPT-6 Sol (max)
88.1 vs 88.0
Grok 4.7 (high) vs GPT-6 Sol (max)
87.8 vs 88.0
Kimi K3 vs GPT-6 Sol (max)
88.5 vs 88.0
Qwen3.8 Max vs MiniMax M3
83.2 vs 82.7
Claude Fable 5.1 (max) vs GPT-5.6 Sol (ultra)
90.5 vs 88.0
GPT-5.6 Sol (high) vs Grok 4.6 (high)
87.6 vs 85.9
GPT-5.6 Sol (ultra) vs Grok 4.6
88.0 vs 86.3
GPT-5.6 Sol (high) vs Kimi K3 (thinking)
87.6 vs 87.2
Qwen3.7 Max (high) vs Gemini 3.1 Pro (default)
80.8 vs 80.1
← Back to the leaderboard