Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 83.4.
Grok 4.5 at #30 and Qwen3.8 Max at #38, 85.80 overall to 83.44. Only one and a half points, and the two fail in opposite directions, which makes this a real choice rather than a ranking. Grok wins voice 87.07 to 81.93, continuity 83.50 to 78.08, and hooks 87.96 to 83.50. Qwen wins length adherence and it is not close: 92.90 against 78.67, over sixteen points. So Grok writes the better script at the wrong length; Qwen hits the word count with flatter prose. Which you prefer depends on which you would rather fix, and trimming a good draft is usually easier than injecting voice into a correctly-sized one. Grok is also four times cheaper at under four cents against 15, and ten times faster at 37 seconds against 397. On price, speed and voice it is the better pick, with length as the one real caveat.
Pick Grok 4.5 for better writing, much cheaper and much faster, if you can trim to length.
Pick Qwen3.8 Max if hitting an exact word count without editing is the thing you are optimising for.
Blue bars: Grok 4.5. Orange bars: Qwen3.8 Max. Same 0–100 scale; the bold bar wins that metric.
| Grok 4.5 | Qwen3.8 Max | |
|---|---|---|
| Overall / 100 | 85.8 | 83.4 |
| Writing Elo | 2042 | 1866 |
| Run-to-run spread (± overall std) | 2.810 | 5.620 |
| Cost per script (USD) | 0.038 | 0.150 |
| Avg latency (s) | 41.7 | 397.4 |
| Open weights | No | No |
Full scorecards: Grok 4.5 · Qwen3.8 Max. How scoring works: methodology.
← All comparisons