Part of: 27 languages, one by one
Polish: drop one "nie" and "I won't have kids" becomes "I will"
· Zeyu Si
On this page
In a father-son scene, the son says: "Ja nie będę miał dzieci." I'm not going to have children.
Speechmatics wrote "Będę miał dzieci." One missing "nie", and the sentence says the opposite while still reading perfectly well. It was the only flipped negation we found across four Polish clips. That one line is why Speechmatics isn't a reference track for Polish.
The test clips: four films, all three difficulty tiers
- Dzień świra (2002): fast father-son banter, clean studio sound. Our baseline tier, and the clip that separated engines best.
- Psy (1992): a police raid. Fast exchanges, interruptions, swearing, background noise. Mid-to-high tier.
- Kiler (1997): a rapid-fire interrogation packed with proper nouns like gun models and city names. Mid-to-high tier.
- Wesele (2004): a wedding party falling apart. Drunk guests shouting over each other, a phone ringing, constant noise. Our hardest tier.
All four are standard Polish, no regional dialect.
Engines we tested
The first round had four: ElevenLabs Scribe v2, AssemblyAI, Speechmatics and Doubao (pl-PL). Soniox transcripts came from a separate all-engine run and were reviewed later. Gemini was added in August 2026. Gladia and Deepgram had already failed on other European languages, so we didn't retest them here.
Red lines, engine by engine
Speechmatics: besides the flipped negation, it inserted three words that nobody said in the wedding scene, "Myślisz że to" ("you think that").
AssemblyAI: no invented sentences, no flips, but it does invent single words. "Zjemy" ("we'll eat") appeared in the raid scene, "otestował" in the wedding scene, and in the clean clip it added "lewicką" ("left-wing") to a line about a Levi's jacket.
Doubao: in the clean clip, during a quick-fire exchange about Chopin pieces, it slipped in "chwalony" ("praised"). Every engine struggled on that stretch. On the first two clips Doubao initially scored 111% and 118% and produced English gibberish. That was our mistake: the wrong language code sent it back to Chinese mode. With the right code it produced coherent Polish.
A spot where everyone failed: in the wedding scene someone says "Psy go zabrali" ("the cops took him"). Psy is slang for police. ElevenLabs and Doubao both heard "Wszyscy" ("everyone"); AssemblyAI wrote "Psego". When several engines fail at the same point, the audio itself is the problem, so we don't hold it against any one of them.
Even in the wedding chaos, none of the four invented a whole sentence.
Error rates on the clean clip: ElevenLabs 5.7%, AssemblyAI 7.8%, Doubao 10.2%, Speechmatics 10.6%. ElevenLabs made one real mistake in the whole clip, writing "Iść" for "Idź".
What we chose, and why
Primary: ElevenLabs Scribe v2. Cleanest across all four clips, no flips, near-perfect on the clean one.
Reference 1: Soniox. Re-checked against our zero-red-line standard, it had no invented words, no flips and no dropped lines across all four clips. By the same measure AssemblyAI had three or more weak inventions and dropped three lines. In the wedding scene Soniox was best in the field. It was the only engine to catch "Co jest?" ("what's going on?") and the only one to get the plural "z nimi" ("with them") right. Doubao wrote "z nią" ("with her").
Reference 2: Doubao. No flips, and it's built differently from the others, so its mistakes tend not to line up with theirs. The negation flip is the main reason it got this slot over Speechmatics.
Reference 3: Gemini. In the August round it had no red lines in Polish. No hallucinations or flips in the raid scene, with the swearing intact. In the wedding scene, only small shifts in word order and person, and it transcribed more of what was actually said than the subtitles did.
AssemblyAI is now first in line as a backup.
Evidence against our choice
- Soniox didn't win on the first pass. In our first review in early July we concluded it was "accurate on the clean clip but not better than AssemblyAI and Doubao, and a tie isn't a reason to switch." A few days later we recounted with a stricter red-line standard, AssemblyAI's weak inventions counted against it, and the decision flipped. A different yardstick can give a different order.
- By error rate, AssemblyAI came first on the interrogation clip (28.2%).
- The primary engine makes mistakes too: ElevenLabs missed the "Psy" slang as well.
Still open
- All four clips are films. No real research interviews.
- No dialects tested.
- Proper nouns (gun models, Chopin titles) trip up nearly every engine. They have to be fixed by the terminology step that runs after transcription.
Written by Zeyu Si