What gets measured, who does the judging, how the ranking is built, and how far you should trust it.
ToneBench asks one question, and I want an honest answer to it: can a model write a full YouTube script the way I write mine? Not a nice paragraph. A whole script, with the hook, the structure, the voice and the technical depth. The current rank compares 167 model configurations that each wrote the same 10 real scripts, five times each, and every draft was scored blind against my finished version.
A script can be grammatically perfect and still fall flat on camera. So the score is split into nine dimensions, each from 0 to 100, and blended into one overall with the editorial weights below. Voice and substance carry the most, because a script that sounds like everyone else, or that says nothing, fails no matter how clean it reads. The small ones, like anti-slop and cue quality, are the tie-breakers between a good draft and one that could air as is.
Seven of the nine are the plain average of the three judges. Length is pure code: it compares the spoken word count to the length of my reference, and running long costs more than running short. The slop score is half code, half judge; see its row in the table.
| Tone & Voice Match | 19% | How precisely it captures the reference's anti-hype conversational voice, first-person 'I was there / I use this' experience, running editorial reactions, and 'told not reported' register. |
| Writing Craft & Clarity | 13% | Sentence rhythm, clarity, paragraph polish, read-aloud quality, and whether the script feels naturally written rather than awkward, stiff, or mechanically assembled. |
| Substance, Accuracy & Value | 15% | Technical depth, factual accuracy, source grounding against the brief/research/reference, and value density: every section should teach or advance something without padding. |
| Continuity & Emotion | 14% | Each paragraph causes the next; a felt emotional register in every section (no Wikipedia blocks); the viewer is transported, not informed. Every section must advance with new information; value-less repetition, restating earlier points in new words, and rehashing recap conclusions are penalized. |
| YouTube Best Practices | 12% | Genuine (non-formulaic) subscribe CTA in the intro, real comment CTA + one forward-looking question at the end, open loop planted and resolved, re-engagement beats, correct intro line, never 'like and subscribe'. CTAs and beats must be genuinely bridged into the content, not a bolted-on checklist. |
| Hook Strength | 10% | Power of the opening 15-30 seconds: surprising / contradictory / second-person or personal-frustration hook, clear value promise, no generic intro. |
| Length Adherence | 8% | Spoken word count against that script's target length; running long is penalized harder than running short. |
| Slop Score (EQ-Bench + ours) | 5% | Combined anti-slop: a 50/50 blend of the judged anti-slop (banned filler/structures, spoken em-dash overuse/style compliance, numbers reframed and pushed to [SHOW:]) and the computed EQ-Bench + our-rules lexical slop (sam-paech/slop-score pipeline + our banned-phrase scan). 100 = no slop. |
| Visual Cue Quality | 4% | Presence and usefulness of inline [SHOW:] cues at every meaningful visual change; spoken text still reads cleanly without them. |
We keep the exact rubric wording private, but these descriptions are what each metric actually rewards.
The benchmark's unit is a task, and each task is built from the script behind one of my "What's AI" videos. It has three parts: a brief describing the video, a research packet with the sources, and my finished script, which is the human reference every draft is judged against.
I use real scripts because synthetic prompts are easy to write and just as easy to overfit. A real brief comes with the messy constraints of an actual video: a target length, one clear call to action, a story that has to land in a specific order. The scripts also stress different muscles on purpose. A model that nails a tight explainer can still fall apart on a vulnerable personal story or a long technical walkthrough.
Every model gets the same prompt, about 11,000 tokens in total: a style guide built from my voice and storytelling guides, the target length, and the brief with its research packet. It never sees my reference script. That's the point: it has to get there on its own.
Generation is open-capability. If the environment a model natively runs in offers web search or tools, it can use them to write the best script it can, which is how you'd actually use it. What it can't do is peek at the answers. Models that run through a command-line tool start in an empty folder with no access to our files, so the reference scripts stay out of reach.
Each model writes every task five separate times, and all five drafts get judged. One lucky draft shouldn't make a model, and five runs is also what lets us measure how much a model varies from one try to the next. Output limits start generous and only grow when a provider reports that a draft was cut off or that the whole budget went to reasoning. A draft that gets cut at its limit is never stored or scored. Every effort or reasoning setting is its own row, so Claude Opus 5.5 (high) and Claude Opus 5.5 (max effort) are compared as the separate configurations they are.
The dashed gold ★ marker on the charts, and the top row of the table, is the human baseline: my finished reference scripts, scored under the exact same rubric and by the same three-judge panel as every model. It's a real bar that a real channel works to. Right now 2 configurations score at or above it, and the best one clears it by 0.9 points. Read the caveat just below before you take that as a model out-writing me. Clicking the human baseline anywhere on the site brings you here.
Here's the caveat. When the judges score a model's draft, the rubric tells them to treat my reference as the top anchor. The reference itself is scored in absolute mode, with nothing to compare it against, and on its own it lands in the low 90s, not a literal 100. So the two numbers come from slightly different setups. A model at or above the line writes at that level on this rubric. That's impressive, but it doesn't mean it beat my script line for line.
The scripts themselves weren't a solo effort. Before one became a reference, six people on our team went through it: two AI engineers, a professional writer, two YouTube reviewers and a fact checker. Then our content lead did a final editing pass. We worked on each script together instead of scoring it separately, so there's no inter-rater reliability score to report, and I'd rather say that than invent one. The line you see is the team-reviewed standard the channel actually ships.
Nobody can score thousands of drafts by hand consistently, so the scoring is done by LLM judges. The obvious worry is that one judge's taste quietly becomes the benchmark. The fix is a panel of three off-the-shelf LLM judges from three model families, one each from Anthropic, OpenAI and DeepSeek, and none of them fine-tuned.
Each judge reads the brief, the research packet, my reference script and the candidate draft. It scores every draft independently with the identical fixed rubric, blind to which model wrote it, and it never sees the word count, since length is measured in code. Before it can score a metric, it has to quote short evidence from the draft. That sounds like a small detail, but it's the main defense against halo: without it, a script that charms the judge on tone can get a warm glow on everything else too. The quotes are stored with every run, so any score can be audited.
For the seven judged metrics, the published score is the plain average of the three judges, and the judged half of the slop score works the same way. When any judge or the rubric changes, the entire board is re-scored in one batch instead of mixing old and new regimes.
Three families, because we measured what one judge does. When we re-judged the same stored outputs with judges from different families, every family scored its own models highest, by different amounts. With three family-disjoint judges, no model is ever scored by its own family alone, and any single judge's stylistic tilt gets diluted by the two that don't share it. That reduces the bias. It doesn't prove the judging is unbiased. Every number in the table below is computed from the stored artifacts and regenerates with the board.
| Scoring judge 1 of 3 | DeepSeek V4.1 Flash (OpenRouter) | The DeepSeek seat, run through OpenRouter. DeepSeek V4.1 Flash took this seat on 2026-09-25, when V4 Flash's own provider stopped serving it, and every draft was re-judged on the new seat in one batch. Its average across the board is 78.15, and its per-model overalls follow the consensus ranking at Spearman ρ=0.9906. |
| Scoring judge 2 of 3 | GPT-5.6 Sol, medium reasoning (Codex CLI) | The OpenAI seat: GPT-5.6 Sol at medium reasoning, run through the Codex CLI on the OpenAI subscription. The OpenAI API is not used for judging. Its average across the board is 81.67, and its per-model overalls follow the consensus ranking at Spearman ρ=0.9787. |
| Scoring judge 3 of 3 | Claude Opus 5, high effort (Anthropic) | The Anthropic seat: Opus 5 at high effort. It runs only through Claude Code, so every Opus 5 score on the board comes from the same route. Opus 5 took this seat over from Opus 4.8 when 4.8 reached end-of-life. Its average across the board is 75.4, and its per-model overalls follow the consensus ranking at Spearman ρ=0.9954. |
| How the three become one score | per-draft, per-metric plain average | On a typical model's overall, the most and least generous judge sit about 6 points apart, yet each seat's ranking tracks the consensus closely (the ρ values above). They disagree on how generous to be much more than on who writes better. Averaging keeps any one seat's taste from setting the bar, and my reference scripts go through the same panel. |
| Blind head-to-head check | two scripts side by side, both orders | A second way of judging, as a check: Opus 5 reads two anonymous scripts side by side, in both orders, without seeing their rubric scores, and picks the better one. It covers neighboring pairs in the top 15 across all 10 current tasks (540 judgments). The two orders agree 68.5% of the time; when they disagree on near-tied neighbors, that pair counts as a draw. Most of these neighbors are near-ties the board doesn't claim to order, so the Spearman of 0.189 against the board is a noisy signal from a sparse fit. The number that could actually contradict the board is this one: on the 6 pairs whose board intervals don't overlap, it agrees on the order 83.3% of the time. That's a descriptive comparison, not a test of each pair's Elo gap. |
One rule sounds boring and matters a lot: every draft on the board needs all three seats. Otherwise a model scored by two judges would sit right next to models scored by three, and you'd have no way to tell.
So judge availability is a gate. Before any new generation or judging starts, the pipeline makes one tiny live call to each seat, through that seat's own route and settings, and checks that the model actually served is exactly the pinned one. If a seat can't run, new work stops. There are only two ways forward: bring that exact judge back, or replace the seat properly. Replacing it means a new seat, every draft on the board re-judged on it, the reference scores re-anchored, and the whole board republished in one batch.
This already happened once, with the DeepSeek seat in September (see its card above).
Every row carries two numbers, and they answer different questions. The overall score is absolute: the weighted blend of the nine metrics, averaged over a model's drafts. Elo is relative: where a model stands against everyone else on this board. Read the overall as "how good is this writing" and Elo as "who beats whom here".
The Elo is built from those same rubric scores. For each of the 10 tasks, every pair of models is compared, and the higher weighted overall wins. There's no separate side-by-side judgment in this step. A win also has to be real: the gap on that task must be bigger than what the two models' own run-to-run noise can explain, using a two-sided 95% Welch t threshold on their five run scores with a one-point floor. Otherwise it's a draw. So "draw" means statistically indistinguishable on that task, not some arbitrary fixed band.
Those wins, losses and draws then go through a standard Elo update (ratings start at 1500, and the K-factor, the step size of each update, is 16), averaged over many shuffled passes so the order of the games doesn't tilt the result.
Every Elo also comes with a 95% interval, our estimate of how much that one model's Elo could move. To get it, we repeatedly resample from every model's recorded run scores on each task and recalculate the full ranking, a couple of hundred times. The interval is the spread you get.
Now, how to read them. When two ranges sit far apart, that's as clear as this board gets: strong evidence for the order, still not proof. When they overlap a lot, the pair is too close to call on this board, and that isn't the same as a tie. Keep the limits of that shortcut in mind, too. Comparing these separate ranges does not test the Elo gap between two models directly. The tasks stay fixed, so the ranges do not measure how the ranking would change on new tasks. Counting pairs by interval overlap describes the board; it doesn't establish pairwise wins or ties.
Now, a fair objection to classic Elo is that it depends on the order you feed it the games, and on the K-factor. Shuffling helps, but it doesn't fully answer the objection. So we also refit the exact same wins, losses and draws with a Bradley-Terry maximum-likelihood model, the method behind LMArena's Chatbot Arena leaderboard, with draws counted as half a win for each side. It has no update order and no K-factor at all, so if the two disagreed, the Elo would be the one to worry about. They don't: Spearman 0.9999 against the published board, the same #1, no model moves more than 4 places, and nobody in the top 10 moves at all.
The weights are an editorial call from the Towards AI team: how much each criterion matters for an AI education video that's useful and that people actually watch through. You'll notice nothing here tries to predict click-through. Clicks depend heavily on the title, the thumbnail, the audience and distribution, and none of that is in the script. What's scored is the script itself: clarity, usefulness, structure, tone, the hook, retention potential, and whether the viewer walks away with something.
They're configurable, and they may move as we get more human feedback, retention data or editorial evidence. If you'd weigh things differently, the weight perturbation check just below shows how much the ranking actually depends on our choice.
A single score hides a lot of ways to be wrong, so a few standing checks are computed with the board and stored in its published data:
These are bounded checks. They can catch a ranking that hinges on one thing; they can't guarantee there's no bias left.
Neither one touches the score, but both matter the moment you have to pick a model for real work. Cost per script is what one script would cost on the provider's API at list price, so every model is compared on the same basis. Models we run on a subscription, like Claude and GPT, don't come with a per-call bill, so we price them as if they were on the API: the provider's own output count for the run, thinking included, plus the fixed prompt billed once. An asterisk on the board means the price itself is an estimate. That makes it a comparison number, not our invoice.
Time is the average time to get one accepted script, retries included. Both are logged on every run and shown on each model's scorecard, and the thinking-level page puts them next to the score for every reasoning-effort setting, so you can see when thinking harder is worth paying for and when it isn't.
A ranking is only fair if everyone ran the same race, so a configuration needs judged drafts for all 10 tasks to hold a current rank.
Sometimes a model can't finish a task on its exact route. The route disappears, or the model keeps failing to return a usable complete script under the same frozen request settings everyone else got. When that happens, the gap is attested and documented. We never fill it with a zero or an estimate, and we never swap in a different model. That configuration comes off the current rank, and its archived scorecard stays visible with the reason. Complete coverage is necessary, and it isn't always enough: some retired routes keep a historical scorecard without a current rank. If any configuration is affected right now, the summary at the top of this page says so.
Contamination, plainly: if a model has already seen my script, it can recall instead of write. New scripts are normally benchmarked before their video goes out. We also withhold non-public references, keep the reference titles and video links off the site, and stamp the benchmark data with an industry-standard canary marker so anyone building a training corpus can filter it out. All of that lowers the risk. None of it makes exposure impossible.
LLM judges scale, but they are still models grading models, and a benchmark about writing for people needs a human check next to them. So we built a blind A/B arena. You see two passages, the same section of the same script brief written by two different models, with no names attached, and you pick the better-written one.
Each passage is a whole script section, somewhere between about 20 and 360 spoken words, and both sides come from the same content beat, so you're comparing two takes on the same material. Passages show as plain prose, without markdown symbols or blank-line gaps. Human scoring also reports a length-control sensitivity: how much passage length alone sways the votes.
Two things to be clear about. It isn't public yet: calibration with my team is in progress behind a login. And human ratings are a separate diagnostic that is never blended into the canonical judge Elo, so no vote in the arena moves a model on this leaderboard. The arena also runs on its own fixed cohort of contestants, which is not the same list as the automated board.
Read this board as directional rather than statistically definitive, especially while the reference set is this small. It compares scripts against Towards AI's editorial criteria for educational videos, and it doesn't try to measure every dimension of how content performs.
To be concrete about what this benchmark does not tell you:
One narrow skill: can a model write a full YouTube script in my voice, at the standard my channel ships? That covers voice match, writing craft, substance grounded in a research packet, YouTube structure and retention, the hook, flow and emotion, length discipline, anti-slop, and on-screen cue quality. It says nothing about general intelligence or coding. If you just want the ranking, that's the leaderboard.
Because they answer a different question. Public leaderboards measure broad preference across everything, which tells you very little about whether a model can hold one voice for two thousand words, follow a style guide, stay away from AI clichés, and land on a target length. ToneBench uses the scripts behind real videos on my channel and normally benchmarks each one before its video goes out. That lowers the contamination risk. It can't prove zero exposure, and we don't pretend it does.
Yes, and we measured it. Under the earlier single-judge setup, we re-judged the same stored outputs with judges from different families, and every family scored its own models highest, by different amounts. That's the reason scoring now goes through a panel of three judges from three model families: Claude Opus 5 (Anthropic), GPT-5.6 Sol (OpenAI), and DeepSeek V4.1 Flash, all scoring every draft blind with the identical rubric. No model is ever scored by its own family alone. A panel dilutes the tilt. It doesn't prove the tilt is gone.
They're the scripts behind ten of my "What's AI" YouTube videos. AI helped with parts of the drafting, but the final versions are mostly human-written and edited, and the Towards AI team reviewed each one before it became a reference. More about the team is on the About page.
If a model has already seen the answer, a writing benchmark turns into a memory test. So new scripts normally join the board before their video goes out, we withhold what isn't public, and we stamp the data with a canary marker. That lowers the risk. It can't rule exposure out.
No. The blind human arena is a separate diagnostic. It's in calibration with my team behind a login right now, and it isn't public yet. Human ratings are never blended into the judge Elo, so no vote there moves a model on this board.