Towards AITowards AIToneBench

Every Alibaba (Qwen) model, side by side

Qwen is a fun family to look at, because the Flash model wins. Qwen3.8 Flash is the best Qwen configuration at #79, and Qwen3.8 Max, the bigger and pricier sibling, lands right behind it. Go back one generation and Qwen3.7 Max is already 220 Elo behind, and the open Qwen3 rows sit far down the board. So with Qwen, the generation you pick matters more than the size.

Best writer
Qwen3.8 Flash#79 · Elo 1676 · $0.013 per script
Cheapest
Qwen3.5 9B#172 · Elo 229 · $0.0031 per script

All 10 configurations

Every configuration here wrote the same 10 scripts for the same four blind judges, so the writing Elo and ranks compare across the whole board, not just inside this family. A highlighted row is a model's best configuration, and the rows under it are the same model at other effort or thinking settings. Scores run 0 to 100, and cost is one full script at the API list price of the route we ran.

Board rankConfigurationEloOverallToneCraftSubstanceHookYouTubeFlowSlopLengthCuesCost / scriptTime / script
#79Qwen3.8 Flash1676 ±4783.282.282.583.383.382.279.488.492.579.7$0.0135.7 min
#81Qwen3.8 Max1645 ±5282.979.881.484.282.581.978.791.194.083.8$0.1617.2 min
#103Qwen3.7 Max (default)1456 ±3780.479.080.279.784.079.678.885.283.478.3$0.0941.8 min
#106Qwen3.7 Max (high)1447 ±4080.278.780.079.183.980.378.485.282.378.1$0.0921.7 min
#143Qwen3-Max1046 ±3474.476.179.271.682.170.675.565.069.869.7$0.0270.7 min
#153Qwen3 235B A22B · open851 ±6870.474.975.668.179.563.369.563.668.155.1$0.00550.8 min
#155Qwen3 235B Thinking · open786 ±4769.970.573.760.181.967.966.063.379.569.4$0.0121.8 min
#166Qwen3 Next 80B · open553 ±4264.969.569.758.980.550.863.864.076.036.6$0.00490.4 min
#172Qwen3.5 9B · open229 ±6357.459.050.956.868.753.150.675.366.840.7$0.00311.2 min
#175Qwen3 8B · open118 ±4255.155.460.458.867.847.551.567.625.169.8$0.00621.0 min

Which one should you use?

This one's easy: Qwen3.8 Flash is both the best writer and the value pick, at $0.013 per script. Qwen3.8 Max scores about the same and costs 12.3 times as much. The only reason to pay that is if length and visual cues matter most to you, since it scores a bit higher on both. One catch with Flash: it's slow, about 5.7 minutes per accepted script, and Max is even slower. If you need open weights, the best open Qwen is down at #153, and Qwen3.5 9B, the cheapest at $0.0031, is too far behind for me to use on a real script.

Thinking levels

Only Qwen3.7 Max ran at more than one reasoning-effort setting here. Its page shows what each step up buys.

Every row links to its full scorecard, and the cross-lab match-ups live in the head-to-head comparisons.

← All comparisons