Head-to-head on the Towards AI writing benchmark: same 10 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.6 (high) leads overall, 88.1 to 83.3.
Grok 4.6 wins this one cleanly. It sits seventeenth on the board at 2144.5 Elo; Qwen3.8 Max is #42 at 1798.7, a gap of 345.8 points, and the confidence intervals do not overlap, so the gap is real. The overall scores say the same thing: 88.14 against 83.35, a difference of 4.79 points. The contrast is sharpest where scripts live or die. Grok hooks at 90.03 to Qwen's 83.95, and holds continuity and emotion at 86.27 where Qwen drops to 78.21. Qwen's one clear win is length adherence, 90.63 to 88.26, and it edges the slop metric 91.96 to 91.84. Consistency matters too: Qwen's overall swings with a 5.43 standard deviation against Grok's 1.43, so its average hides rougher outings. Then the unusual part: the better model is also the cheaper one. Grok runs about twelve cents per script, Qwen about fifteen cents, both exact figures. Neither is open weights, so there is no philosophical tiebreaker. This one is not close.
Pick Grok 4.6 (high) if you want the stronger script on nearly every metric, especially hooks and continuity, and you would rather pay less for it.
Pick Qwen3.8 Max if strict length adherence is the metric you care about most and you can accept wider swings in overall quality.
Blue bars: Grok 4.6 (high). Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| Grok 4.6 (high) | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 88.1 | 83.3 |
| Writing Elo | 2144 | 1799 |
| Run-to-run spread (± overall std) | 1.430 | 5.430 |
| Cost per script (USD) | 0.115 | 0.150 |
| Avg latency (s) | 199.2 | 395.6 |
| Open weights | No | No |
Full scorecards: Grok 4.6 (high) · Qwen3.8 Max. How scoring works: methodology.
← All comparisons