Guide
How to choose a transcription tool: measure it on your audio, then price your hours
A four-step way to pick transcription software, an API or a human service: decide what kind of service you need, expect accuracy to depend on your recordings more than on the brand, test two or three candidates on ten minutes of your own audio with a word error rate calculator, and price your real monthly hours. With measured numbers and sourced prices.
No paid links on this page: vendor links go straight to the vendor. Affiliate policy.
Most comparisons of transcription tools rank them by features and the vendor's own accuracy claim. Two things matter more: how accurate a tool is on recordings like yours, and what your monthly hours cost on each plan. Both can be measured in an afternoon.
1. Decide what kind of service you need
| You need | Choose | Typical price |
|---|---|---|
| Searchable notes of meetings or interviews, edited by you | A transcription app (Otter, Rev, Sonix, Happy Scribe, Trint) | $16.99 to $100 a month |
| To edit a podcast or video by editing its transcript | An editor built on transcription (Descript, Riverside) | $24.00 to $35.00 a month for the first paid plan |
| Text in your own software, or hundreds of hours | A speech-to-text API (OpenAI, Deepgram, AssemblyAI, Google, AWS, Azure) | $0.15 to $0.96 an audio hour |
| Every word right: legal, medical, publication | Human transcription (Rev, Happy Scribe) | $119 to $120 an audio hour |
Prices are US list prices read on 7 October 2026; the cost calculator has every plan, its limits and its source.
2. Expect accuracy to depend on the recording
In the baseline benchmark, the same open-source model, Whisper medium.en, missed or mangled:
| Recording | Words wrong |
|---|---|
| Clean read speech (audiobook) | 1.9% |
| The same speech with restaurant noise | 3.7% |
| A relaxed two-person video call, speakers with different first languages | 13.6% |
| A four-person meeting, all headsets mixed | 22.6% |
A vendor's "up to 99% accurate" usually describes the first row. If your recordings look like the last two rows, expect several times as many errors from any tool, and test before you commit. In the meetings, most errors were words left out entirely: quick exchanges, asides, and the ends of turns where someone else starts talking.
3. Test two or three tools on ten minutes of your own audio
- Pick a representative recording, the kind you will transcribe most often, about ten minutes long. Use a recording you are allowed to upload: check that the participants agreed to it, and read each tool's privacy terms on how long uploads are kept and whether they are used to train its models.
- Make a reference. Take the best transcript you get and correct it word by word while listening. Expect it to take several times the length of the audio. It is the step that makes the test worth anything.
- Run each candidate on the same file with default settings, using its free trial.
- Score each transcript in the WER calculator with the Standard preset, so that punctuation, "25" against "twenty-five" and filler words are not counted as errors.
- Read the alignment, not only the rate. A tool that drops "not" is worse than one that misspells a surname, even at the same WER. Names, numbers and technical terms are worth checking one by one.
Ten minutes is about 1,500 words. A difference of 1 to 2 points between tools on one file is within the noise; a difference of 5 points is real. If two tools are close, choose on price and features.
4. Price your real monthly hours
Enter your hours, your longest recording and your billing preference in the cost calculator. Three things the plan pages make easy to miss:
- Allowances and per-file limits. Otter Pro includes 20 hours a month but at most 10 imported files and 90 minutes per recording. Descript's allowance counts every file you upload, transcribed or not. Happy Scribe's plans count uploaded files separately from live meetings.
- Overage. Sonix charges $10.00 an hour past any plan's allowance and Happy Scribe $12.00; several others publish no overage price, so the next plan up is the only option.
- Yearly billing lowers the monthly price, from 8% on Sonix Core to 51% on Otter Pro, but you pay for twelve months up front.
At ten hours a month, the APIs cost a few dollars, the apps $15 to $100, and human transcription about $1,200. The apps are not overpriced APIs: the editor, speaker labels, search, exports and support are what you pay for. If you only need text and can run a script, an API is the cheapest route by far; the open-source Whisper models in the benchmark cost nothing but your own computer's time.
Other things to check
- Speaker labels. Most apps label speakers; on some APIs (AssemblyAI and Deepgram among them) it is an add-on charged per minute.
- Languages and accents. If your speakers have accents or switch languages, test with them; a single clean sample tells you little.
- Exports. Subtitles need SRT or WebVTT with sensible line lengths; the subtitle converter converts between formats and fixes line length and timing in your browser.
- Data handling. Where the audio is stored, for how long, and whether it trains the vendor's models. This is in the privacy policy or data processing terms, not the pricing page.
- Terms. Some tools forbid commercial use on free plans, or automated access. Read them if you will use the tool for client work.
How this guide was made
The accuracy figures come from the baseline benchmark: open-source models run on openly licensed recordings with human transcripts, scored by the same code as the WER calculator. Prices come from each vendor's own page on the date shown. Commercial tools' accuracy is not compared here because none could be tested within our rules yet (methodology); no tool is recommended over another on accuracy until it has been measured. Links to Descript and Riverside may become affiliate links, labeled "(paid link)" when they do; that does not change what is measured or how it is described (affiliate disclosure).
Also available as Markdown.