Free tool
Word error rate (WER) calculator: alignment, substitutions, deletions and insertions
Paste a correct transcript and a machine transcript to get the word error rate and character error rate, the substitutions, deletions and insertions, and a word-by-word alignment showing every error. Normalization options are documented one by one. The same code scores the TranscribeBench benchmark, and it matches jiwer's counts.
Word error rate (WER) calculator
Runs in your browser. Nothing you paste is sent anywhere.
The formula
WER = (S + D + I) / N
S substitutions: a reference word replaced by a different word
D deletions: a reference word missing from the transcript
I insertions: a word in the transcript that is not in the reference
N words in the reference (N = S + D + correct words)
The three counts come from the alignment that turns the reference into the transcript with the fewest word edits (the Levenshtein distance, computed on words). When several alignments are equally short, the calculator picks the same one as jiwer, the Python library most ASR papers and leaderboards use, so its S, D and I match jiwer's as well as its WER. Its test suite checks this on 349 jiwer-generated cases, including transcripts of up to 400 words.
A worked example. The reference is "the cat sat on the mat" (6 words) and the transcript is "the cat sit on mat". "sat" became "sit" (1 substitution) and the second "the" is missing (1 deletion). WER = (1 + 1 + 0) / 6 = 33.3%.
WER can exceed 100%. Insertions have no ceiling: "hello" transcribed as "hello there my friend" is 3 insertions over 1 reference word, a WER of 300%. A WER of 0% means every word matched after normalization.
Character error rate (CER) is the same formula over characters of the normalized texts, with the single space between words counted as a character, as jiwer does. It is kinder to near-misses: "colour" for "color" is one word error but one character error out of six.
Several clips. Across many recordings, add up the errors and the reference words, then divide: 3 errors in a 10-word clip and 0 in a 90-word clip is 3% over the whole set, not the 15% you would get by averaging the two clips' rates. The benchmark reports rates this way.
Normalization, option by option
A transcript can be right about every word and still score badly because it writes "Mr." where the reference says "mister", or "25" where it says "twenty-five". Normalization removes differences of writing style before the words are compared. It is applied to both texts in the same way, and you can see the normalized texts under the result.
| Option | What it does | Example: these count as equal |
|---|---|---|
| Ignore case | Lowercases everything. | "Paris" = "paris" |
| Ignore punctuation | Everything except letters, digits and apostrophes inside words becomes a space. Hyphens split words. A period or comma between digits is kept (3.5), and a comma between digit groups is dropped (1,000 = 1000). | "well-known." = "well known" |
| Numbers as digits | Number words become digits: cardinals with "and" ("one hundred and five" = 105), years said in pairs ("nineteen ninety-nine" = 1999, "nineteen oh five" = 1905, from thirteen up, so the clock time "ten thirty" stays 10 30), decimals ("three point one four" = 3.14), ordinals ("twenty first" = 21st, but a lone "second" stays a word). "$5" = "5 dollars", "5%" = "5 percent", "£" and "€" likewise. | "twenty-five dollars" = "$25.00" |
| Expand contractions | won't = will not, can't and cannot = can not, n't = not, 're = are, 've = have, 'll = will, 'd = would, I'm = I am, let's = let us, y'all = you all, gonna = going to, wanna = want to, gotta = got to, ain't = is not. "'s" becomes "is" only after it, he, she, that, there, here, what, who, where, when, why, how, this and a few similar words; after a name it is usually possessive and is left alone. | "we're" = "we are" |
| Ignore filler words | Removes um, umm, uh, uhh, uhm, er, erm, hm, hmm, mm, mmm, mhm. "Uh-huh" and "mm-hmm" become "huh" and nothing, since they split at the hyphen. | "um so we" = "so we" |
| Ignore bracketed tags | Removes anything in [square], <angle>, {curly} or (round) brackets, such as [laughter], <overlap> or (inaudible). | "I [laughs] know" = "I know" |
| Spell out titles | Mr = mister, Mrs = missus, Ms = miss, Dr = doctor (before a capitalized name or with a period). | "Mr. Smith" = "mister Smith" |
The Standard preset turns all of them on. It is the setting the benchmark uses, written down before any results were collected. Basic ignores only case and punctuation. None compares the texts exactly as typed, which is right when the exact form matters (legal transcripts, subtitles with house style).
What normalization does not do: it does not treat British and American spellings as equal ("colour" and "color" differ), does not match "OK" with "okay", and does not judge whether two different words mean the same thing. Those count as errors, as they would in most published WER figures. If one matters for your use, fix it in the reference or read the alignment.
Using WER to choose a tool
- Use your own audio. WER depends on the recording far more than on the tool: the same model can score 3% on a clean audiobook and 30% on a noisy meeting. The benchmark shows how large that spread is.
- Ten minutes is a start, an hour is better. A few hundred words give a rough number; small differences between tools need more.
- Check the reference. Many "errors" turn out to be mistakes in a hurried reference transcript. Look at the alignment before trusting the number.
- WER is not everything. It ignores punctuation, capitalization (with normalization on), speaker labels and timestamps, and it treats a wrong name the same as a wrong "the". For subtitles or quotes, read the errors, not only the rate.
Privacy and the code
Both texts stay in your browser. The scorer is a small JavaScript module, core.js, the same file the benchmark runs on its transcripts, so any figure on this site can be reproduced by pasting the published reference and transcript into the calculator.
Also available as Markdown.