ToneBench AI Writing Benchmark & Model Leaderboard
Towards AI logo Towards AI ToneBench

The Towards AI Writing Benchmark · LLM Leaderboard

ToneBench

Can a model write in our voice?

Every model writes the same real YouTube scripts, and a panel of three LLM judges from three different families scores each draft blind against our own finished versions across nine writing dimensions. Read the methodology for what we measure, the human baseline, the judges and trust checks, and why the reference scripts stay private, or the about page for where this comes from.

3-judge panel scored — models — tasks multiple runs, averaged updated —

Key findings

Writing Elo

ToneBench · by Towards AI
drag / scroll →
The best configuration of each model under the selected metric, ranked; every effort level and thinking variant stays in the table below. Drag or scroll horizontally. The dotted-gold ★ bar is our human editorial baseline. Top five tinted.

Writing quality vs price

ToneBench · by Towards AI
One dot per model (its best configuration for the selected metric), colored by family. Click a dot to pin its logo and name. Dashed lines are the medians, splitting the board into four quadrants; up-left is the value frontier. Log-scale cost; free models pinned at the left edge.
← scroll the table sideways to see every metric →
Each cell is a 0–100 score, colored by strength (green good, amber okay, red weak). Hover a column header for what that metric means. Click any model row to open its full scorecard, per-article breakdown, and write-up.
Cost is what one task would cost on the provider's API at list price, so every model is comparable. * means the price itself is an estimate. ~ means the token counts were computed from the prompt and the finished draft rather than reported per call: that applies to the models we run on a subscription (Claude, GPT), and it understates models that spend most of their output on hidden reasoning.
How this is scored

The 9 scripts, the nine metrics and what each rewards, the human baseline, the judge and cross-family trust checks, and why the references stay private. It is all on the methodology page.

Read the methodology →

Explore each metric

the full board matrix is above · here is each dimension on its own, side by side

Slide through the nine writing dimensions. Each card ranks every model on that one metric, so you can see who wins on voice, who wins on hooks, who keeps it tight. Drag, scroll, or use the arrows. Click a model to open its scorecard.

Frequently asked questions

Quick answers about what ToneBench measures, who is behind it, and how far to trust the numbers. The full setup lives on the methodology page; the story and the team are on the about page.

What is ToneBench?

ToneBench is a writing benchmark that asks one narrow, practical question: can an AI model write a YouTube script in our editorial voice? Every model writes the same set of real video briefs, the same articles our team actually produced, and each script is scored against our finished reference version on voice, writing craft, substance, structure, hooks, flow, length discipline, and AI-writing "slop". It is a measure of voice and craft fidelity, not general intelligence.

Who is behind ToneBench?

It is built and maintained by Louis-François Bouchard, co-founder & CTO of Towards AI and creator of the "What's AI" YouTube channel, together with the Towards AI editorial team. More on the team and where this comes from is on the about page.

Why build a custom benchmark instead of using public leaderboards?

Public leaderboards measure broad preference: which answer a crowd or a judge likes better, on average, across everything. That is useful, but it says little about whether a model can hold a specific voice, follow a strict style guide, resist AI clichés, and hit a target length on a real production brief. ToneBench measures exactly that gap. And because the tasks are our own unpublished briefs, models can't have memorized the answers.

How are models scored?

Every model receives the same inputs: our writing style guide, a task brief, and a research packet. Each model writes every script several times (multiple runs, averaged), and every script is scored blind by a panel of three LLM judges from three different families (Anthropic, OpenAI, DeepSeek) against a fixed rubric. Each judge must quote evidence from the script before scoring each metric, every judge uses the identical rubric and prompt for every model, and the published per-metric score is the unweighted mean of the three judgments. Per-metric scores are combined with editorial weights, and the Elo ranking comes from pairwise comparisons of those scores with bootstrap confidence intervals. The exact rubric wording and reference scripts are withheld to keep the benchmark hard to game. The full metric list, weights, judge details, and verification checks live on the methodology page.

Aren't LLM judges biased toward their own family?

Yes. We measured it, and it is exactly why the published score comes from a panel of three judges from three different families: Claude Opus 5 (Anthropic), GPT-5.6 Sol (OpenAI), and DeepSeek V4 Flash. When we re-judged the same stored outputs with different families as judge, every family scored its own models highest, by different amounts. In the panel, each judge scores every draft blind with the identical rubric, the published score is the unweighted mean of the three, and no model is ever scored by its own family alone. We also run blind side-by-side A/B checks in both orders and an order-independent Bradley-Terry refit. The current per-judge agreement numbers live on the methodology page.

What do the metrics mean?

Tone: does it sound like our voice and persona. Craft: sentence-level writing quality. Value: substance and factual grounding in the provided research. Hook: strength of the opening. YouTube: structure and retention practices for the format. Flow: continuity and emotional pacing. Slop: absence of over-used AI phrases and patterns. Length: discipline against the target spoken length. Cues: quality of the visual [SHOW:] directions.

What is the human baseline?

The dotted-gold line is our own team-reviewed reference scripts scored under the same rubric (how the human baseline is made). Those scripts went through multiple expert reviewers and a final editorial pass, so they represent the quality bar the channel actually ships: a high bar, not a theoretical ceiling. The current model-versus-baseline relationship is regenerated in the key findings and board data.

Why are the reference scripts and briefs kept private?

To prevent contamination. If the exact briefs, research packets, and gold scripts were public, future models could train on them and the benchmark would quietly measure memorization instead of writing ability. New articles are benchmarked before their videos are published, the addition date is recorded, and all benchmark data carries an industry-standard canary marker.

How often is the board updated?

New models are added as soon as practical after release. We watch the major provider APIs daily. New articles are added a few times per year, benchmarked pre-publication. When the rubric or judge version changes, the entire board is re-judged under the new setup so scores are never mixed across judging regimes; the evaluation version is recorded with every result.

Model A is one place above model B. Is it actually better?

Maybe. Every model carries a 95% confidence interval; when two intervals overlap, the board does not claim their order is settled. Treat them as a tie. Many widely separated models are distinguishable, but inside the tightly-packed top ten many neighbours are genuine near-ties. The interval, not the rank number, is the honest read.

Does the #1 model here mean it's the best LLM overall?

No. ToneBench measures one demanding, real-world skill: writing long-form scripts in a specific human voice under a strict style contract. That correlates with careful instruction-following and style control, but it is deliberately narrow: a model can top this board and still be mid-pack at code, math, or agentic work. Use it as one signal, alongside benchmarks for the skills you care about.

Can I suggest a model, report an issue, or use the benchmark?

Yes. Missing models, wrong numbers, methodology disagreements, or interest in a similar evaluation for your own brand voice: send feedback. If your team wants help choosing or deploying writing models in production, that is exactly what Towards AI does.