Skip to main content

Part of: 27 languages, one by one

Norwegian: "we weren't unfaithful" became "we chose to be unfaithful"

· Zeyu Si

On this page

In The Worst Person in the World (Verdens verste menneske) there's a party scene where someone says "Vi var ikke utro": we weren't unfaithful. AssemblyAI wrote "Vi valgte utro": we chose to be unfaithful. The negation vanished, the verb changed, and the meaning flipped. The line comes up again in the scene, and AssemblyAI got it wrong every time. Three times in all.

Noise didn't cause this. The scene is at a normal pace, without sustained overlapping speech.

The test clips

Five film clips from three films:

Clip Scene Difficulty
King's Choice, part 1 A warship's guns firing, artillery commands over the noise Extreme
King's Choice, part 2 An air raid on a train Extreme
Headhunters, part 1 An emotional two-person scene, relatively clear Medium-high
Headhunters, part 2 24 seconds of panicked shouting in a car Extreme
The Worst Person in the World A party: normal pace, frequent changes of speaker, relatively clear Medium-high

All five are films. Two are relatively clear, but they're still scripted drama. So when we assessed Gemini in August, we also took a five-minute excerpt from a Norwegian podcast.

Engines we tested

ElevenLabs Scribe v2, AssemblyAI, Speechmatics, Soniox, Gemini, Qwen, FunASR and Doubao. We also looked at NB-Whisper, the open-source model from the National Library of Norway, but didn't adopt it.

Doubao doesn't support Norwegian. We tried all three Norwegian language codes and got English gibberish or nothing at all. FunASR produced only about half as many words as the reference.

The red lines

Negation flips matter most here. Besides the opening line, AssemblyAI turned "I'm willing to do anything" into "I won't do it", and "you must never quit" into "I must…". Speechmatics flips too, in the other direction: it tends to add an ikke, so "know" becomes "don't know". No more than three times across the five clips. And at one negation, every engine in the comparison got it wrong except ElevenLabs.

Invented names: AssemblyAI opened the clear emotional scene with a character called Bård, and put someone called Martha into the air raid. Neither character exists in those scenes. Qwen invented a surname, Hågen.

Inserted content: in the warship clip, Gemini made up several commands that aren't in the film, such as "Gun 1 in position" and "Torpedo battery 3/4".

Language drift: once, Gemini returned a Norwegian clip as a Chinese translation ("What's your name? Julia."). Run again on the same audio, it produced normal Norwegian.

Loops: Gemini repeated itself in the air-raid clip. Its WER there was 89.6%, against 41.8% to 47.8% for the other three engines, a gap of 42 points.

What we chose

  • Primary engine: ElevenLabs Scribe v2. No hallucinations, no flips, no invented names across the five clips. Its mistakes are mishearings: it wrote the opening line as "Vi var ikke ute", we weren't out, but kept the negation.
  • Reference engines: Soniox, Speechmatics, Gemini.

AssemblyAI is out. Four of the five clips had flips or invented names, including the clear ones.

Soniox came in during July. Until then, Speechmatics was the only reference engine for Norwegian. Clip for clip, Soniox beat Speechmatics four times and tied once, and the two mostly went wrong in different places. Soniox got the opening line right, and the name Julie.

Speechmatics never invented a sentence or a name in the five clips, and it also got the opening line right.

Gemini, on the five-minute podcast: no hallucinations, flips or loops. Three sound-alike errors, and two places where it was more accurate than ElevenLabs.

Our Norwegian configuration carries one extra rule: where the engines disagree about an ikke, the transcript marks it as uncertain for a person to check.

Evidence against our choices

  • Gemini crossed a red line on every one of the five film clips: flips, invented commands, loops, and the Chinese output. Our case for running it rests on one five-minute podcast with no reference transcript, checked only against ElevenLabs. In July we had rejected it.
  • Soniox and Speechmatics made three identical directional errors in the emotional scene. The "they fail in different places" argument doesn't always hold where it matters most.
  • An early automated scan found Soniox producing 1.5 times as many words as the reference in Norwegian.

What's still open

  • We don't have a single real conversation with a reference transcript. The evidence for Gemini is especially thin.
  • Negation is still the biggest risk in Norwegian. A case like that one negation only ElevenLabs got right can't be rescued by voting. For now, we mark it as uncertain and let a person listen.

Our Norwegian transcription page

Written by Zeyu Si