Towards AITowards AIToneBench

Qwen3.7 Max (high) vs Gemini 3.1 Pro (default)

Head-to-head on the Towards AI writing benchmark: same 9 real YouTube scripts, five runs each, scored blind by a three-family judge panel. Qwen3.7 Max (high) leads overall, 81.5 to 80.2.

Qwen3.7 Max (high)
#55 Elo 1674 · 81.5/100
Gemini 3.1 Pro (default)
#63 Elo 1577 · 80.2/100
Cost / script
$0.054 vs $0.118
Human baseline
92.7 Qwen3.7 Max (high) falls below it · Gemini 3.1 Pro (default) falls below it

The verdict

Qwen3.7 Max wins this one outright now: its Elo floor (1628) sits above Gemini's ceiling (1614), so the sliver of doubt I used to flag here is gone. What isn't in doubt is the one metric with a real gap: length adherence, 83.3 for Qwen against Gemini's 75.9. Gemini keeps missing the target script length, and in a video pipeline that has a genuine cost; a draft that runs long means cutting on the timeline or re-prompting and waiting again. On voice they're basically the same writer, 80.4 vs 80.0, and both are clearly short of sounding like me. Gemini claws some back on visual cues, 81.5 against Qwen's 81.0, one of Qwen's weakest metrics, and it returns drafts about two and a half times faster. Cost favors Qwen at roughly 5 cents per script, estimated, versus Gemini's 12, though neither number will hurt at this scale. If I were automating one of these two into a weekly pipeline tomorrow, Qwen's length discipline is what settles it.

Pick Qwen3.7 Max if script length is what your pipeline keeps fighting and you'd like to pay less per draft.
Pick Gemini 3.1 Pro if you need faster drafts and better visual cues and can fix script length yourself.

Metric by metric

Blue bars: Qwen3.7 Max (high). Orange bars: Gemini 3.1 Pro (default). Same 0–100 scale; the bold bar wins that metric.

Tone & Voice Match19% weight
80.9
79.9
Writing Craft & Clarity13% weight
81.5
80.8
Substance, Accuracy & Value15% weight
81.2
80.8
Continuity & Emotion14% weight
79.2
78.4
YouTube Best Practices12% weight
80.7
78.9
Hook Strength10% weight
85.0
82.3
Length Adherence8% weight
83.3
75.9
Slop Score (EQ-Bench + ours)5% weight
85.9
87.7
Visual Cue Quality4% weight
78.4
81.7

Everything else that differs

Qwen3.7 Max (high)Gemini 3.1 Pro (default)
Overall / 10081.580.2
Writing Elo16741577
Run-to-run spread (± overall std)2.5802.420
Cost per script (USD)0.0540.118
Avg latency (s)157.263.8
Open weightsNoNo

Full scorecards: Qwen3.7 Max (high) · Gemini 3.1 Pro (default). How scoring works: methodology.

← All comparisons