Skip to main content

Field notes from 27 languages: how we picked the engines for each

· Zeyu Si

On this page

We support 27 languages. The engines behind each one were tested separately. There's no single configuration applied across the board.

This series opens up the testing for every language: which clips we used, which engines we tried, what mistakes each one made, what we chose and why, and what's still unresolved.

The short version

There is no "best engine". An engine that leads in Spanish can come last in Japanese. One that makes no mistakes on clean audio can drop whole passages once there's a dialect or people talking over each other.

For each language we look for the combination that makes the fewest mistakes in that language and fails in different places from the others. The second half matters as much as the first: engines that keep failing on the same words can't correct each other. We explain why in Why one engine isn't enough.

The testing method itself is in WER lies: real conversation as the reference, check the reference first, then count errors that change the meaning.

The current line-up

Each language has one main engine, which produces the transcript and tells speakers apart, plus two or three supporting engines that cross-check it.

The main engine is currently ElevenLabs Scribe v2 for every language. The supporting engines fall into eight combinations:

Supporting engines Languages
AssemblyAI + Doubao English, German, Spanish, French, Italian, Dutch, Portuguese, Russian, Turkish, Ukrainian
Soniox + Doubao Indonesian, Japanese, Malay, Polish, Romanian, Thai
Soniox + Speechmatics Arabic, Czech, Norwegian
Speechmatics + Soniox Danish, Greek
AssemblyAI + Soniox Hungarian, Swedish
Doubao + Soniox Korean, Vietnamese
Speechmatics + AssemblyAI Finnish
Doubao + FunASR + iFlytek Chinese (the only language with three)

Since August 2026, Gemini has also been running as an additional supporting engine for all 27 languages.

Notes for each language

Languages we don't support yet

We tested Tamil, Bengali and Filipino but couldn't find a usable reference for any of them. Tamil subtitles are conventionally written in the literary register while people speak the colloquial one. For Bengali we couldn't find verbatim human subtitles. And some programmes labelled as Filipino turned out to be in a different language.

If we can't measure it, we don't offer it. Once we have a reliable reference, for example a verbatim transcript from a native speaker, we'll test again.

Written by Zeyu Si