Skip to main content

The Transcribe Labs Blog

How we test transcription quality, and what we found.

Start here

By topic

All posts

  1. Why one engine isn't enough: more engines only help if they fail differently

    One engine gets names and numbers consistently wrong. More engines only help when they fail in different places. A real interview, including where we lost.

    Continue reading
  2. The most dangerous transcription errors read perfectly

    A flipped negation reads perfectly and a majority vote won't catch it. So tools should show where they're unsure. We checked 15 of them.

    Continue reading
  3. WER lies. Here's how we measure accuracy on real interviews

    The industry's favourite error metric misleads on interview audio. What we use instead: check the reference first, then count errors that change meaning.

    Continue reading
  4. When interviews switch languages

    Spanish interviewee, Chinese researcher, an interpreter in between. Vendor docs and lab tests misled us; a real recording gave the opposite answer.

    Continue reading
  5. Who said that? Speaker attribution when many people talk

    Built-in speaker labels drift as people are added. Learning each voice first got us above 99% on public meeting data, as long as people take turns.

    Continue reading
  6. Field notes from 27 languages: how we picked the engines for each

    No engine is best everywhere. The index to our language-by-language tests: which engines we picked for each of 27 languages, and why.

    Continue reading

Subscribe via Atom feed