Skip to main content

Part of: 27 languages, one by one

Korean: three engines, one identical wrong negative

· Zeyu Si

On this page

One line in our Korean tests came out wrong in exactly the same way from three engines. There's no negative in what the speaker says, yet Doubao, AssemblyAI and Soniox all wrote 쓰지 않았어, a sentence with "didn't" in it.

When three engines make the same mistake, the audio at that spot is the problem, not any single engine. If we'd dropped Soniox for it, we'd have had to drop the two engines already in use as well. It was also a reminder that a three-way vote doesn't protect you on negatives.

The test clips

Four Korean film scenes, scored against Korean subtitles for the deaf and hard of hearing:

  • Memories of Murder, an interrogation: three people, with laughter, groaning and overlapping speech.
  • Memories of Murder, the final interrogation: four or five people, shouting and swearing.
  • Extreme Job, the scene where they name the fried chicken: several people talking fast over each other, plus cheering.
  • The Accidental Detective 2, a client briefing: four people in rapid back-and-forth, with a hacking demo in the middle.

The first two are top level on our three-level difficulty scale and the other two are middle level. There's no easy clip. The first two come from official scene clips released by the studio.

Engines we tried

ElevenLabs Scribe v2, AssemblyAI, Doubao and Speechmatics in the first round. Soniox, Gladia and others in follow-up runs. Gemini was tested separately.

Where the engines crossed red lines

Flipped meaning. Besides the shared error above, Speechmatics had a reversal of its own in the detective scene, where 오빠 넣으면 and the negative 안 오면 ("if he doesn't come") got swapped.

Inserted content. In the chicken scene Speechmatics invented a name, 박동규야. Doubao's weak spot is numbers: it produced "5 존맛" out of nowhere and wrote a stretch of crying as "6 6".

Repetition loops. In the Memories of Murder laughter, Doubao wrote the laugh as ㅎ repeated 40 times. It never collapsed like that on the two middle-level clips; it only happened in extreme laughing and shouting. Gladia also looped on Korean. Soniox was the only engine that didn't collapse on either of the two hardest clips.

Dropped passages. AssemblyAI's biggest problem is leaving things out. It skips whole chunks and often returns the fewest characters of the four. It rarely invents anything, but it loses a lot.

What we chose

Primary: ElevenLabs Scribe v2. First on all four clips, with a mean of 11.6% against 17.0% for the runner-up. It's also the only one that regularly gets names right.

References, in order: Doubao, Soniox, Gemini. All three run on every Korean file.

  • Doubao: second on average, and a steady second on both middle-level clips. Its collapses only happened in extreme laughter, which rarely comes up in a real interview.
  • Soniox: it replaced AssemblyAI, and not because of scores. We found that the worst kind of AssemblyAI error was one Soniox tended to make too, often in the same spot, while the errors unique to AssemblyAI were its large dropped chunks. Swapping in Soniox leaves the reference tracks with less overlap in what they get wrong.
  • Gemini: runs as the third reference.

Evidence against our choice

  • Our first four-clip Korean report recommended something else: Doubao plus AssemblyAI, on the grounds that AssemblyAI "only drops, never invents, so it's safest". Review then showed AssemblyAI was one of the three engines behind the shared negative, so "safest" didn't hold up, and Soniox took its place.
  • Gemini's Korean record is weak. When we tested it, Korean ended up in the "don't add" group. In the chicken scene it changed the meaning four or five times, turning "the juices are still in" into "absolving the flesh", and left content out. In the detective scene it changed the meaning three times, turning "thought it was a body" into "thought he was asleep", and dropped a "5%". Most of these are meaning-changing substitutions, which aren't among our six red lines, but they're just as dangerous in an interview transcript. It's in the configuration now, and it needs watching.
  • Doubao's phantom numbers are the kind of thing a reader could easily take as data the participant gave.

Open problems

  • The two hardest clips were only checked for error rate and collapse at the time, not line by line for meaning.
  • All four clips are films. We have no real interview. The configuration still carries a note saying the whole pipeline should be run end to end on a real interview before full production use.
  • A three-way shared error can't be fixed by voting. So when engines disagree on whether something is negative, our merge step doesn't go with the majority; without hard evidence it flags the line for a person.

Our Korean transcription page

Written by Zeyu Si