Rules

  1. Measured means measured. No accuracy figure appears on this site unless it was computed here, by the code described below, on audio with a human reference transcript. Vendor claims ("99% accurate") are not repeated as facts.
  2. Rules before results. The test set, the selection rules, the normalization and the scoring were fixed before any model or tool was run, and are published with the results.
  3. Everything reproducible. The clip list, the scripts, every reference transcript and every transcript we scored are published. The audio comes from public corpora and is rebuilt byte for byte by a script.
  4. Only licensed audio. Every clip comes from a corpus whose license allows this use. Only those clips are ever uploaded to a tool.
  5. Terms first, free tiers only. A commercial tool is tested only through a free tier that needs no account or card, and only if its terms allow benchmarking, publishing results and automated access. Vendors see their results seven days before publication.
  6. Unknown stays unknown. A price or limit a vendor does not publish is shown as "Check vendor", never estimated.

The test set

23 clips, 33.6 minutes, 5,242 reference words, in five conditions. The list with every clip's exact source is benchmark-samples.yaml.

Condition Clips Source and license How the clips were chosen
Clean read speech 5 LibriSpeech test-clean, CC BY 4.0 3 women and 2 men drawn with a fixed random seed from the corpus's speaker list. Each clip is the speaker's first chapter, consecutive utterances in order, stopping before 75 seconds.
Harder read speech 4 LibriSpeech test-other, CC BY 4.0 Same rule, 2 women and 2 men. The corpus authors put speakers that an early recognizer found harder into this split.
Restaurant noise 5 The five clean clips, plus "Restaurant ambience" (public domain) Each clean clip mixed with the noise at a 5 dB signal-to-noise ratio, from a random point in the recording (fixed seed).
Accented conversation 5 EdAcc test set, CC BY-SA 4.0 Conversations in corpus order, taking one when a speaker's first language is not English and no non-English first language repeats one already taken, until five. In each, the first window at least 2 minutes in, 90-150 seconds long, that starts and ends where nobody is speaking and has no read-aloud passage.
Meetings 4 AMI Meeting Corpus test meetings ES2004a, IS1009a, TS3003a, EN2002a (headset mix), CC BY 4.0 In each meeting, the first window at least 5 minutes in, 120-180 seconds long, that starts and ends where nobody is speaking.

Signal-to-noise ratio. SNR = 10 log10(speech power / noise power), with power the mean square of the samples over the whole clip, pauses included. At 5 dB the speech is about three times as powerful as the noise (10^0.5 = 3.2).

Reference transcripts. Each comes from the corpus itself and was not edited:

When people talk over each other, a single transcript has to put one person's words first. The references order speech by segment start time, so a tool that orders overlapping speech differently is charged for it. That is one reason meeting error rates are high everywhere.

Privacy. Speakers are known only by the corpora's anonymous IDs. For EdAcc we read only the first languages each speaker reported, and nothing else from its survey. We never try to identify anyone, and we do not publish the audio.

Not in this pilot. Mozilla Common Voice (CC0) is now distributed only through the Mozilla Data Collective, which requires an account; whether to register is the site owner's decision. Public-domain government recordings with official transcripts, telephone audio and non-English speech are planned.

Word error rate, plainly

Word error rate (WER) counts the fewest single-word edits that turn the correct transcript (the reference) into the transcript being scored, divided by the number of words in the reference:

WER = (S + D + I) / N
S = substitutions (a wrong word), D = deletions (a missed word), I = insertions (an extra word)
N = words in the reference

A WER of 5% means about one word in twenty is wrong, missing or extra. Lower is better; 0% is perfect, and WER can go above 100% when a transcript adds many words. The edits come from a word-level Levenshtein alignment; where several alignments are equally short, the same one as the widely used jiwer library is chosen, so the S, D and I counts match jiwer's as well as the total.

Character error rate (CER) is the same calculation over characters, including the spaces between words. It gives partial credit for near-misses ("realise" for "realize" is one character wrong, not one word).

Over several clips the errors and the reference words are added up before dividing, so a long clip counts for more than a short one. The interval shown with each overall figure is a 95% bootstrap interval: the clips are resampled with replacement 2,000 times (fixed seed) and the middle 95% of the resulting rates is reported. It shows how much the figure could move with a different handful of clips from the same sources. It does not cover differences in recording conditions outside the test set.

The code is the WER calculator itself: the benchmark sends every reference and transcript through the same JavaScript module, then re-scores every clip with jiwer as an independent check (no difference on any clip).

Normalization

Before words are compared, both texts are normalized the same way, so that a transcript is not charged for writing style: case, punctuation, "25" against "twenty-five", "we're" against "we are", filler words such as "um", bracketed tags such as [laughter], and "Mr." against "mister". The benchmark uses the calculator's Standard preset, which applies all seven rules; each is described with examples on the WER calculator page.

Normalization does not equate British and American spellings, "OK" and "okay", or words that mean the same thing. Those count as errors, as they do in most published figures. A different normalization gives different numbers, which is why figures from different sources are hard to compare: The Whisper paper, for example, scores with OpenAI's own English normalizer, whose rules differ from these.

Open-source baselines

Baselines are open models run on our own machine, so they can be measured without anyone's permission and re-run by anyone. They show what the test set's conditions do to accuracy and give a yardstick for commercial tools.

Commercial tools

Before any use, each tool's terms of service and acceptable-use policy are read and the relevant clauses recorded. A tool is tested only if:

  1. it has a free tier usable through its normal web page without an account and without a payment card, and
  2. its terms do not forbid benchmarking, publishing results, competitive analysis or automated access.

Runs use a scripted browser that does not disguise itself, one file at a time at a human pace, with default settings, uploading only the licensed clips. No account is created (accounts in the site owner's own name are his decision), and no captcha, bot check, quota, rate limit or paywall is worked around: where a tool stops us, we stop and record it. Every transcript is scored by the same code as the baselines. Before publication, each vendor receives its own results and has seven days to reply; replies are published alongside, and results change only if a vendor shows an error in the method.

In the October 2026 pilot, no commercial tool met both conditions with a complete transcript, so no commercial result is published. The tool-by-tool record will be published with the first commercial results.

Prices

The cost calculator uses list prices read from each vendor's own pricing page or official documentation, for the US, before tax, with the date each was read and the quote it came from (prices.json). Promotions are noted, not used. Where a page did not state a number, it is null and shown as "Check vendor". Where a vendor's own pages disagreed, the note says which figure was used and why. Prices are re-read monthly and whenever a vendor announces a change.

Limits

Conflicts of interest

TranscribeBench may earn affiliate commissions from some vendors (affiliate disclosure; none are active today). That never decides which tools are tested, how they are scored or the order of results.

Corrections

Corrections are dated and described on the page they affect. If a price, a clip, a reference or a score here is wrong, the contact details are on the about page.

Also available as Markdown.