Towards AITowards AIToneBench

How ToneBench works

What we measure, how we score it, the human baseline, and how we track reference exposure.

ToneBench asks one question and tries to answer it honestly: can a model write a full YouTube script in our voice? The current rank compares 130 models that each wrote the same 9 real scripts, five times each, with every draft scored blind against our own finished versions. We keep 130 configurations visible.

What we measure

One demanding, real-world skill: writing long-form scripts in a specific human voice under a strict style contract. Nine dimensions, each scored 0 to 100 and blended into the overall by the editorial weight below. Voice and substance carry the most; the smaller signals like anti-slop and cue quality are the tie-breakers that separate a good draft from one that ships.

Tone & Voice Match19%How precisely it captures the reference's anti-hype conversational voice, first-person 'I was there / I use this' experience, running editorial reactions, and 'told not reported' register.
Writing Craft & Clarity13%Sentence rhythm, clarity, paragraph polish, read-aloud quality, and whether the script feels naturally written rather than awkward, stiff, or mechanically assembled.
Substance, Accuracy & Value15%Technical depth, factual accuracy, source grounding against the brief/research/reference, and value density: every section should teach or advance something without padding.
Continuity & Emotion14%Each paragraph causes the next; a felt emotional register in every section (no Wikipedia blocks); the viewer is transported, not informed. Every section must advance with new information; value-less repetition, restating earlier points in new words, and rehashing recap conclusions are penalized.
YouTube Best Practices12%Genuine (non-formulaic) subscribe CTA in the intro, real comment CTA + one forward-looking question at the end, open loop planted and resolved, re-engagement beats, correct intro line, never 'like and subscribe'. CTAs and beats must be genuinely bridged into the content, not a bolted-on checklist.
Hook Strength10%Power of the opening 15-30 seconds: surprising / contradictory / second-person or personal-frustration hook, clear value promise, no generic intro.
Length Adherence8%Spoken word count vs target; overshooting is penalized harder than undershooting (aim for the short end, ~3 Google-Docs pages).
Slop Score (EQ-Bench + ours)5%Combined anti-slop: a 50/50 blend of the judged anti-slop (banned filler/structures, spoken em-dash overuse/style compliance, numbers reframed and pushed to [SHOW:]) and the computed EQ-Bench + our-rules lexical slop (sam-paech/slop-score pipeline + our banned-phrase scan). 100 = no slop.
Visual Cue Quality4%Presence and usefulness of inline [SHOW:] cues at every meaningful visual change; spoken text still reads cleanly without them.

We keep the exact rubric wording private, but this is genuinely what each metric rewards. Anti-slop, for example, is a 50/50 blend of a judged pass (banned filler and AI-signature structures, spoken em-dashes) and a lexical slop metric adapted from EQ-Bench.

The 9 scripts

Real briefs from the "What's AI" channel, not synthetic prompts. They stress different muscles: a model that nails a tight explainer can still fall apart on a vulnerable personal story. The benchmark does not republish non-public references, and source-publication timing is disclosed when a task was already public.

Article 1
opinion / warning explainer
A share-this-now opinion piece: something the author had to get off his chest, arguing a clear position with receipts.
~1546 spoken words
Article 2
news-analysis / skeptical explainer
A viral-moment react: explain the new buzzword plainly, bring the receipts that it isn't new, and land what genuinely changed.
~1257 spoken words
Article 3
personal roadmap / opinion
A first-person roadmap essay. Heavy on lived experience, here-is-how-I-would-do-it, and honest trade-offs.
~1777 spoken words
Article 4
short explainer
A tight teach-one-concept video in the RAG/CAG cadence: relatable problem, plain-English mechanism, one good analogy, warm sign-off.
~2080 spoken words
Article 5
news-analysis / opinion explainer
A deep, fair news breakdown: explain from first principles, react to it, acknowledge the good before the criticism.
~1678 spoken words
Article 6
founder announcement / personal origin story
The most personal register: a held-back reveal, a vulnerable origin story, real dates as narrative, gratitude at the close.
~1175 spoken words
Article 7
personal technical walkthrough / agentic workflow case study
A first-person technical case study: trace a real video-production workflow from manual bottlenecks through agent-assisted research, writing, packaging, and human review.
~4260 spoken words
Article 8
product release announcement / personal observation
A real production brief from the channel, evaluated under the same fixed style contract.
~876 spoken words
Article 9
career guide / hiring analysis
A real production brief from the channel, evaluated under the same fixed style contract.
~3085 spoken words

The human baseline & editorial review

The dotted-gold ★ line on the charts, and the top row of the table, is our own human baseline: the finished reference scripts scored under the exact same rubric as the models. It is a real bar, not a theoretical ceiling. No model reaches it on the current board. Clicking the human baseline anywhere on the site brings you here.

The reference scripts went through a collaborative editorial process by the Towards AI content/editorial team. This included review from 2 AI engineers, 1 professional writer, 2 YouTube/script-optimization reviewers with experience in YouTube content, and 1 fact reviewer with strong writing and editing expertise. A final editing pass was then completed by the Towards AI content lead. This was not a blind independent annotation setup, so no inter-rater reliability score is reported. Instead, the team worked collaboratively toward a final version that satisfied the editorial criteria. The human baseline represents this team-reviewed editorial standard rather than an average of isolated reviewer scores.

So the human number is not one reviewer's opinion. It is the standard the channel actually ships, run through the same three-judge panel and rubric every model faces, which is what makes it a fair top line to measure against.

The judges, and how we keep them honest

Every published score is the consensus of three off-the-shelf LLM judges from three different model families: Anthropic, OpenAI, and DeepSeek. Each judge scores every draft independently with the identical fixed rubric, blind to which model wrote it, and must quote short evidence from the draft before scoring each metric, so metrics are judged on their own signal instead of a halo. The published per-metric score is the unweighted mean of the three judgments. When any judge or the rubric changes, we re-score the entire board in one batch rather than mixing regimes.

Why a panel: when we re-judged the same stored outputs with judges from different families, every family scored its own models highest, by different amounts. Averaging three family-disjoint judges means no model is ever scored by its own family alone, and any single judge's stylistic tilt is diluted by the two that don't share it. Every number below is computed from the stored artifacts and regenerates with the board.

Scoring judge 1 of 3Claude Opus 5, high effort (Anthropic)Canonical judge: Opus 5 at high effort through Claude Code subscription only. New direct Anthropic API or Message Batches evidence is prohibited; historical transport artifacts remain audit records only. Opus 5 succeeded Opus 4.8 when 4.8 reached end-of-life. Board-mean 71.43; its per-model overalls track the published consensus at Spearman ρ=0.995. Fixed rubric, evidence quotes required before every metric score, blind to model identity, never sees word counts.
Scoring judge 2 of 3GPT-5.6 Sol, medium reasoning (Codex CLI)Runs through the Codex CLI at medium reasoning effort. Board-mean 79.79; its per-model overalls track the published consensus at Spearman ρ=0.9893. Fixed rubric, evidence quotes required before every metric score, blind to model identity, never sees word counts.
Scoring judge 3 of 3DeepSeek V4 Flash (OpenRouter)Runs through OpenRouter; the panel's fastest, cheapest voice. Board-mean 78.75; its per-model overalls track the published consensus at Spearman ρ=0.9766. Fixed rubric, evidence quotes required before every metric score, blind to model identity, never sees word counts.
How the three become one scoreper-draft, per-metric unweighted meanEvery draft is scored independently by all three judges with the identical rubric prompt; the published per-metric score is their unweighted mean. Mean spread between the highest and lowest judge on a model's overall is 9.45 points. No family's models are ever scored by their own family alone, and the human-baseline reference scripts go through the same panel.
Blind head-to-head A/B protocolclaude-code:claude-opus-5Current Opus 5 different-protocol check (adaptive thinking): anonymized side-by-side verdicts in both orders on adjacent top-15 pairs across all 9 current tasks (486 judgments). Order consistency is 63%; disagreements between near-tied neighbours are counted as draws. Transport provenance is mixed and preserved: 324 earlier verdicts are frozen historical direct-API evidence, while the 162 task-8 verdicts used Claude Code subscription. Every future refresh is subscription-only. Its sparse head-to-head fit has Spearman 0.254 vs the board, but none of these adjacent pairs has non-overlapping board confidence intervals, so this is a tie-region stress test rather than a settled-rank validation.
Order-independent refitBradley-Terry MLE (deterministic)The exact pairwise outcomes refit with the Chatbot-Arena-style Bradley-Terry model: Spearman 0.9999 vs the published Elo, same #1, max rank shift 3. The ranking is not an artifact of update order or K-factor.

Elo, scores, and confidence

The overall score is the weighted blend of the nine metrics. Elo comes from comparing those overalls pairwise across the 9 scripts over many shuffled passes, with bootstrap confidence intervals on every rating. Read Elo as the relative ranking and the overall as the absolute quality. When two models' intervals overlap, the board is not claiming one beats the other.

As a standing cross-check, the same pairwise outcomes are also fit with a Bradley-Terry maximum-likelihood model (the method behind LMArena's Chatbot Arena leaderboard, with draws counted as half a win for each side). The deterministic BT ranking agrees with the published board at Spearman 0.9999, picks the same #1, and no model moves more than 3 places (top 10: at most 1). The published Elo is not an artifact of update order or the K-factor.

Why the weights are what they are

The specific metric weights are domain-expert editorial weights chosen by the Towards AI team. They reflect how much importance our team gives each criterion for useful, watchable AI educational content. This benchmark is not designed to optimize only for CTR. CTR is heavily affected by title, thumbnail, audience, and distribution, while this benchmark focuses on script quality: clarity, usefulness, structure, tone, hook quality, retention potential, and whether the content brings value to the viewer. These weights are configurable and may evolve as more human feedback, retention data, or editorial evidence becomes available.

Reference privacy and exposure

Contamination, plainly. Public reference material can make a writing benchmark measure recall instead of writing ability. New tasks are normally evaluated before source publication, and every task records its addition and source-publication timing when known. We withhold non-public references, record the dates, and stamp benchmark data with an industry-standard canary marker. The disclosed timing is part of how results should be interpreted, not a claim that exposure is impossible.

Limitations

This benchmark should be interpreted as directional rather than statistically definitive, especially while the reference set is small. It is designed to compare outputs against Towards AI's editorial criteria for educational video scripts, not to measure all possible dimensions of content performance.

Frequently asked questions

What does ToneBench measure, exactly?

One narrow, real skill: can a model write a full YouTube script in our editorial voice? Not general intelligence, not coding. It scores voice-match, writing craft, substance grounded in a research packet, structure and retention for the format, hook strength, flow and emotion, length discipline, anti-slop, and on-screen cue quality. If you want the models ranked, that is the leaderboard.

Why not just use a public leaderboard?

Public leaderboards measure broad average preference across everything. They say little about whether a model can hold one voice, follow a style guide, resist AI cliches, and hit a target length on a real brief. ToneBench uses first-party production briefs, normally evaluates them before publication, and discloses source timing for exceptions. That reduces contamination risk without pretending it can prove zero exposure.

Aren't LLM judges biased toward their own family?

Yes. We measured it. Re-judging the same stored outputs with judges from different families showed every family scoring its own models highest, by different amounts. That is exactly why the published score is a panel of three judges from three different families: Claude Opus 5 (Anthropic), GPT-5.6 Sol (OpenAI), and DeepSeek V4 Flash, all scoring every draft blind with the identical rubric. Each metric's published score is the unweighted mean of the three, so no model is ever scored by its own family alone. Exact judge ids and live agreement numbers are in the table above.

Who builds and reviews the reference scripts?

The Towards AI editorial team, and the writing itself comes from the "What's AI" YouTube channel. More about the team is on the About page.

How does ToneBench handle reference exposure?

Contamination can turn a writing benchmark into a memorization test. We normally evaluate new tasks before source publication, record addition and publication dates, withhold non-public reference material, and stamp benchmark data with a canary marker. New tasks are normally evaluated before source publication, and every task records its addition and source-publication timing when known.

← Back to the leaderboard