Part of: 27 languages, one by one
Dutch: no engine crossed a red line. So how do you choose?
· Zeyu Si
On this page
Dutch was the calmest language in this whole exercise. Four recordings, four engines, and not a single invented sentence, flipped negation or runaway loop. Every mistake was a mishearing at the level of one word: a word swapped for something that sounds like it, where people talked over each other, spoke too fast or had a heavy accent.
When red lines can't tell engines apart, you need another way to choose. We looked at whose mistakes differ from whose.
The test clips: none of them easy
- Alles is Liefde (2007), two scenes: a crowded TV-studio scene with people cutting each other off, plus background noise; and a fast argument over a supermarket PA announcement. Mid-to-high tier.
- OP1 debate on Afghanistan (2021): a TV round table with guests arguing and a host jumping in. The only real interview-style conversation in the Dutch set. Mid-to-high tier.
- New Kids Turbo (2010): a debt collector at the door, Brabant dialect, several people going at it, heavy noise, fast speech. Hardest tier.
None of the four is a clean, steady-paced baseline clip.
Engines we tested
First round: ElevenLabs Scribe v2, AssemblyAI, Speechmatics and Doubao (nl-NL). Soniox transcripts were later reviewed line by line. Gemini was added in August 2026. Gladia and Deepgram had already failed on other European languages and weren't tested for Dutch.
Everyone's mistakes were small
- ElevenLabs heard "Tonnen" as "Donder" in the studio scene, and twice wrote "uitverkocht" (sold out) for "uitverkoop" (clearance sale) in the PA announcement.
- AssemblyAI turned "Sint maakt grapjes" (Sinterklaas is joking) into a sound-alike string, "Zit maar krampje".
- Speechmatics turned "Ik kan die man niet verstaan" (I can't understand that man) into "I can't understand someone".
- Doubao got it right where the subtitles were wrong. The subtitle scan had misread "Pardon" as "Rardon" and "Pleur op" as "Rleur op". Doubao wrote the correct words and was marked wrong for it.
One trap caught almost everyone. In Brabant dialect "you" is "gij", and several engines heard it as "ik" (I). The subject of the sentence changed from "you" to "I". Only AssemblyAI got it right.
Soniox or Doubao: a tie, so which one stays?
In our review Soniox met our bar for a reference track: zero red lines, and first among the engines other than the primary. The reviewer suggested swapping it in for Doubao. So we compared the two line by line across all four clips:
- Studio scene: Soniox was better. Doubao dropped two whole sentences and turned "give him" into "give me".
- Argument: Doubao was better. It got a key line, "Pleur op!" (get lost!), which Soniox garbled into meaningless syllables.
- Interview: identical in meaning. Soniox's lead came entirely from names of people and institutions, and from Doubao filling the page with "Uh".
- Dialect scene: both went wrong at the same spots in the same way.
Close enough in meaning that it wasn't a reason to switch.
What we chose, and why
Primary: ElevenLabs Scribe v2. The most conservative of the four, never invents words. Cleanest on the interview, with every name right.
Reference 1: AssemblyAI. What we worry about most with AssemblyAI is looping and invented sentences. Neither showed up in the interview. And it was the only engine that didn't fall into the "you/I" dialect trap.
Reference 2: Doubao. No red lines on any clip, and it's built differently from the others, so its mistakes tend not to overlap with theirs.
Reference 3: Gemini. No red lines on the interview. It got the key quotes right (roughly, "you have the watches, we have the time"), along with names like NATO and The Hague, and transcribed more of the conversation than any other engine. On the argument clip it made one slip: "I'll come get the sand" became an unrelated word, "Conversant".
Soniox is first in line as a backup, ahead of Speechmatics.
Evidence against our choice
- By our own standard, Soniox qualified, and the reviewer recommended the swap. Keeping Doubao was partly about valuing a different technical approach, not purely about measured quality.
- Gemini's Dutch evidence rests on one conversational clip. That's thin.
- By error rate, AssemblyAI came first on two of the three usable clips. ElevenLabs, our primary, didn't come first outright on any of them.
Still open
- Belgian Dutch (Flemish) hasn't been tested.
- We have only one real interview clip, and its subtitles are a rewrite, so we can judge meaning but can't score it.
Written by Zeyu Si