The Transcribe Labs Blog
How we test transcription quality, and what we found.
Start here
- Why one engine isn't enough: more engines only help if they fail differently
- The most dangerous transcription errors read perfectly
- WER lies. Here's how we measure accuracy on real interviews
- Who said that? Speaker attribution when many people talk
- Field notes from 27 languages: how we picked the engines for each
By topic
How we measure accuracy
Mixed languages
Multiple speakers
27 languages, one by one
All posts
Why one engine isn't enough: more engines only help if they fail differently
One engine gets names and numbers consistently wrong. More engines only help when they fail in different places. A real interview, including where we lost.
Continue readingThe most dangerous transcription errors read perfectly
A flipped negation reads perfectly and a majority vote won't catch it. So tools should show where they're unsure. We checked 15 of them.
Continue readingWER lies. Here's how we measure accuracy on real interviews
The industry's favourite error metric misleads on interview audio. What we use instead: check the reference first, then count errors that change meaning.
Continue readingWhen interviews switch languages
Spanish interviewee, Chinese researcher, an interpreter in between. Vendor docs and lab tests misled us; a real recording gave the opposite answer.
Continue readingWho said that? Speaker attribution when many people talk
Built-in speaker labels drift as people are added. Learning each voice first got us above 99% on public meeting data, as long as people take turns.
Continue readingField notes from 27 languages: how we picked the engines for each
No engine is best everywhere. The index to our language-by-language tests: which engines we picked for each of 27 languages, and why.
Continue reading