WER lies. Here's how we measure accuracy on real interviews
· Zeyu Si
On this page
We once tested a book-review interview from Swedish morning television. Four speech engines transcribed it. By WER, the metric almost everyone uses, their error rates came out between 194% and 216%. More errors than words.
Then we read the four transcripts line by line. They were the most complete of the whole batch. Book titles, names, plot points, all correct. They even caught things the reference transcript had left out.
The problem was the reference.
This post covers two things: why WER breaks down on interviews, and what we switched to. It's the method we used to choose engines for all 27 of our languages.
Five ways WER misleads you on interview audio
We ran into the same few problems across more than 20 languages. Each one produces a number that looks precise and says the wrong thing.
1. The reference has already been trimmed. Most of our references come from subtitles for the deaf and hard of hearing on films and TV. Subtitles have to be readable at speed, so they routinely compress what people actually say. The more faithfully an engine transcribes, the more it disagrees with the subtitle, and the more it gets penalised. That was the Swedish interview. In a Spanish film scene where everyone shouts over everyone else, every engine scored between 63% and 80% WER. Read line by line, every one of them had got the meaning across.
2. The written language and the spoken language aren't the same thing. Written Arabic and the Arabic people speak day to day can be far apart. We scored the same engine against a reference written in Modern Standard Arabic and got 49%. Against a reference that faithfully recorded the dialect, 16%. The engine hadn't changed. The reference had.
Tamil is more extreme. Subtitles are conventionally written in the literary register, while people speak the colloquial one. Use subtitles as the reference and the winner is whichever engine "rewrites" speech into literary Tamil. The engine that transcribes faithfully comes last. The ranking flips.
3. Chinese and Japanese don't put spaces between words. Split by spaces and a whole Chinese sentence becomes one "word". The error rate can come out above 1000%. That number means nothing.
4. Spelling differs, meaning doesn't. Greek and Arabic often score high. Look closely and a lot of the "errors" are inflection and spelling conventions, such as whether a small Arabic mark called a hamza was written. No reader would misunderstand the sentence. The metric counts it anyway.
5. So WER is a smoke test, nothing more. It tells you whether an engine has fallen over completely: whole passages missing, screens of gibberish. For deciding which of two working engines is more accurate, we no longer look at it.
References should be real conversation
Real interviews have pauses, filler words, interruptions and people talking over each other. Test on newsreaders or read-aloud audio and every engine looks great. Interviews are a different story.
So our test clips are spontaneous speech only. Film scenes come first: they have subtitles for the deaf and hard of hearing that track the dialogue closely, and they're available. Where we can't find a suitable film, we fall back in order to TV series and documentaries, parliament or court recordings, then podcasts with verbatim transcripts.
There are three difficulty tiers, each built on the one before:
- Basic: one or two people, steady pace, no background noise, standard accent, but still real conversation. This tier alone knocks out a few weak engines.
- Harder: basic plus one or two complications, such as faster speech, a small group, quicker turn-taking, some background noise or an accent.
- Extreme: overlapping speech, heavy noise or music, strong dialect, poor audio, all at once.
When a language has strong regional variation, each region gets its own clips. Spanish from Spain, Mexico and Argentina. Arabic from the Gulf, Egypt and the Levant.
Check the reference before you check the engine
This is the step people skip. We added it after it cost us. Before a reference is used to score anything, we confirm the reference itself is right.
In one audit we threw out a lot of clips that had looked fine:
- A programme labelled as Tagalog was actually in Cebuano, a different Philippine language.
- A Thai clip had its subtitles about 30 seconds out of sync with the audio.
- A Cantonese interview turned out to be in Korean. The guest was Korean and the Cantonese subtitles were a translation. Every engine scored close to 100% error on it.
Each of these would have produced a precise-looking, meaningless score. Worse, some earlier conclusions had been built on clips like these.
There are three languages we still can't find a usable reference for: Tamil, Bengali and Filipino. We don't support them yet. We'd rather not offer a language than put out a number we can't stand behind.
Count errors that change the meaning
Researchers read interview transcripts for meaning. Writing "um" as "uh" doesn't matter. Writing "didn't" as "did" does. So for each engine we count red-line errors, in six categories:
- Invented sentences: a sentence appears that nobody said. One engine with a perfectly respectable error rate produced, in a clip with a strong accent, a complete sentence ending "buy some bananas". Nobody said anything like it. We didn't use that engine.
- Flipped negation: "no" becomes "yes" or the other way round. It reads perfectly and means the opposite. Our next post is about this one.
- Looping: the same word repeated dozens or hundreds of times.
- Inserted content: material added that wasn't in the audio.
- Dropped passages: a stretch of speech missing entirely. In a film scene with heavy dialect and a noisy crowd, one engine dropped the first 54 seconds.
- Language drift: switching languages partway through. One engine turned long passages of Spanish into Portuguese. "Soy abogado" became "Sou advogado".
Every engine is measured with the same ruler. Differences in how names are spelled, or whether a number is written "6" or "six", don't count as red lines.
People reading transcripts miss things too, so we also run two machine checks:
- Output ratio: the number of words an engine produced divided by the number in the reference. Well below 1 usually means large passages were dropped; we've seen 0.1. Well above 1 with lots of repeated words usually means looping; we've seen 1.8.
- Repeats: how often the same word recurs within a short span.
Neither depends on audio format or clip difficulty, so they can be compared across languages. WER can't.
Rules for not fooling ourselves
Each of these came out of a mistake.
When several engines agree and only the reference disagrees, suspect the reference first. We got this wrong twice in Turkish. We logged real speech as "hallucination" when the subtitles had simply left that line out.
When several engines make the same mistake in the same place, don't pin it on one of them. Usually the audio is genuinely unclear there and anyone would get it wrong. Counting it against one engine is unfair to that engine, and it leads you to pick the wrong one.
Short clips can't be trusted. Always run the full recording. One approach we looked at scored 99.8% on speaker identification over a 10-minute clip. Run on the full recording, it dropped to 63.9%.
Run every setting at least twice. The same input doesn't always give the same output. Once, an engine repeated a single "um" 99 times and swallowed the next three sentences. Run it once and you can't tell whether that's typical or a fluke.
Where this method falls short
Most of our test clips come from film and television, not from research interviews. Film dialogue is scripted and recorded with care. A single recorder on a meeting-room table is a different thing. So what these results tell you is how engines fail on real conversation, not how accurate your particular interview will be.
That's also why we don't publish a single headline accuracy figure for transcription. What we can share is how we test, what we found, and what we still don't know.
Next: the most dangerous of the six red lines, where the meaning flips and the sentence still reads perfectly.
Written by Zeyu Si