Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 leads overall, 86.6 to 81.1.
Grok 4.6 wins this one comfortably. It sits at #26 with 2035.3 Elo against Qwen3.8 Max at #46 and 1713.9, and the confidence intervals don't come close to touching, so the 321.4-point gap is real. Overall it's 86.61 to 81.06, a 5.55-point spread, and the damage lands where a script lives or dies: continuity and emotion, 85.78 vs 76.22, YouTube best practices, 86.54 vs 77.76, and tone and voice, 87.86 vs 80.65. Qwen also swings much harder from script to script, a 7.89 standard deviation against Grok's 2.53. It takes exactly one metric back: hook strength, 83.39 to Grok's 80.81. The records say the same thing, Grok at 941 wins and 99 losses against Qwen's 588 and 258. Normally the loser gets a price argument, but there isn't much of one here: about sixteen cents per script against roughly twenty cents, both closed weights. Four cents doesn't buy back a 5.55-point gap.
Pick Grok 4.6 if you want the clearly stronger script on nearly every metric — voice, continuity, structure — delivered with far more consistency at roughly twenty cents each.
Pick Qwen3.8 Max if you want slightly stronger hooks and the cheaper draft at about sixteen cents per script, and you're prepared to rewrite around a 5.55-point overall gap.
Blue bars: Grok 4.6. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| Grok 4.6 | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 86.6 | 81.1 |
| Writing Elo | 2035 | 1714 |
| Run-to-run spread (± overall std) | 2.530 | 7.890 |
| Cost per script (USD) | 0.204 | 0.164 |
| Avg latency (s) | 243.1 | 429.0 |
| Open weights | No | No |
Full scorecards: Grok 4.6 · Qwen3.8 Max. How scoring works: methodology.
← All comparisons