Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Grok 4.5 leads overall, 85.8 to 81.5.
The rare matchup where one model wins on quality, price, and speed at the same time. Grok 4.5 scores 85.8 overall to Qwen3.7 Max's 81.5, costs about two cents a script against Qwen's estimated five, and returns a draft in under 30 seconds versus over two and a half minutes. The Elo intervals don't overlap either. Where the writing actually differs is tone. Grok sits at 87.1 on voice match, Qwen at 80.7, and that gap is wider than both standard deviations combined, so I treat it as real: Qwen reads assembled, not spoken. Its visual cues are also a weak spot here at 81.0. Credit where due, though. Qwen wins length adherence, 83.3 to Grok's 78.7, and Grok's length variance is no joke, so Qwen drafts come back the right size more often. That matters if trimming is the part you hate. But both are closed models, so Qwen can't play the open-weights card, and paying more for the weaker script is a tough sell. Grok, comfortably. Neither is anywhere near the human scripts yet, for the record.
Pick Grok 4.5 if you want the stronger voice at less than half the cost and a fraction of the wait.
Pick Qwen3.7 Max if hitting target length on the first pass is your biggest pain point.
Blue bars: Grok 4.5. Orange bars: Qwen3.7 Max (high). Same 0–100 scale; the bold bar wins that metric.
| Grok 4.5 | Qwen3.7 Max (high) | |
|---|---|---|
| Overall / 100 | 85.8 | 81.5 |
| Writing Elo | 2042 | 1674 |
| Run-to-run spread (± overall std) | 2.810 | 2.580 |
| Cost per script (USD) | 0.038 | 0.054 |
| Avg latency (s) | 41.7 | 157.2 |
| Open weights | No | No |
Full scorecards: Grok 4.5 · Qwen3.7 Max (high). How scoring works: methodology.
← All comparisons