Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.7 Max (high) leads overall, 81.5 to 80.2.
Qwen3.7 Max wins this one outright now: its Elo floor (1628) sits above Gemini's ceiling (1614), so the sliver of doubt I used to flag here is gone. What isn't in doubt is the one metric with a real gap: length adherence, 83.3 for Qwen against Gemini's 75.9. Gemini keeps missing the target script length, and in a video pipeline that has a genuine cost; a draft that runs long means cutting on the timeline or re-prompting and waiting again. On voice they're basically the same writer, 80.4 vs 80.0, and both are clearly short of sounding like me. Gemini claws some back on visual cues, 81.5 against Qwen's 81.0, one of Qwen's weakest metrics, and it returns drafts about two and a half times faster. Cost favors Qwen at roughly 5 cents per script, estimated, versus Gemini's 12, though neither number will hurt at this scale. If I were automating one of these two into a weekly pipeline tomorrow, Qwen's length discipline is what settles it.
Pick Qwen3.7 Max if script length is what your pipeline keeps fighting and you'd like to pay less per draft.
Pick Gemini 3.1 Pro if you need faster drafts and better visual cues and can fix script length yourself.
Blue bars: Qwen3.7 Max (high). Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.
| Qwen3.7 Max (high) | Gemini 3.1 Pro (default) | |
|---|---|---|
| Overall / 100 | 81.5 | 80.2 |
| Writing Elo | 1674 | 1577 |
| Run-to-run spread (± overall std) | 2.580 | 2.420 |
| Cost per script (USD) | 0.054 | 0.118 |
| Avg latency (s) | 157.2 | 63.8 |
| Open weights | No | No |
Full scorecards: Qwen3.7 Max (high) · Gemini 3.1 Pro (default). How scoring works: methodology.
← All comparisons