Part of: 27 languages, one by one
Malay: the speaker switched to English, and the engine translated it back
· Zeyu Si
On this page
In a Malaysian current-affairs interview, the guest switches to English partway through: "enforcing the letter of the law, without looking at … mood music around you."
ElevenLabs wrote that English down word for word. AssemblyAI produced about fifteen words of Malay instead. Roughly the same meaning, perfectly fluent, and not what the guest said.
For researchers, this is the most dangerous error in Malay. The engine has done a translation on the interviewee's behalf, and nothing on the page gives it away.
The test clips: three, chosen for code-switching
- A scene from a Malay TV drama: a mother and son talking, then a drug-squad arrest with shouting and a few English lines ("under arrest" and the like).
- Two Malaysian political talk-show interviews: commentators switching between Malay and English sentence by sentence, whole English clauses, and a lot of names of people and parties.
There was a fourth clip, a TEDx talk, but its volunteer subtitles were a summary, 170 words for three minutes, so it couldn't serve as a reference and we dropped it. Free, conversational, verbatim Malay material is scarce, and we stopped at three. The clips were picked for how much code-switching they contain, not matched to our three difficulty tiers. We didn't specifically test heavy noise with people talking over each other.
Engines we tested
Eleven: ElevenLabs Scribe v2, Soniox, Doubao (ms-MY), AssemblyAI, FunASR, Qwen, Speechmatics (enhanced and Melia-1), Gladia and Deepgram, plus Gemini, added in August 2026.
Red lines, engine by engine
Rewriting English as Malay: AssemblyAI and Gladia both did it on the legal passage above. In the other interview AssemblyAI also turned "asking for the removal" into the Malay for "restoration", the opposite meaning. Gladia turned "we'll look at the comments" into "we'll suppress the comments".
Flipped meaning:
- Deepgram turned "is being investigated" into "escaped".
- FunASR turned the drama line "you've got the wrong person" into "you've got the right person".
- Speechmatics turned the English "I don't know what that means" into "I know What I means".
Lost passages: Deepgram cut off the drama after 2:11, losing roughly 30% of it. FunASR skipped about 20 seconds of one interview.
Inserted words: Speechmatics slipped "menipu" ("to deceive") into a sentence.
Loops: the Melia-1 version of Speechmatics repeated "eh" about fifty times at the end of the drama, and produced a run of "ada ada ada" in an interview. That run was first blamed on Qwen; the review found we had pinned it on the wrong engine.
What we chose, and why
Primary: ElevenLabs Scribe v2. Cleanest on code-switching across all three clips, English clauses kept word for word, and the most accurate on names.
Reference 1: Soniox. No hallucinations or flips on any clip, and the lowest error rate on all three (only 0.4 points ahead of AssemblyAI on the drama). Its one small slip was writing "latter" for "letter".
Reference 2: Doubao. No hallucinations or flips either, and the only engine that never "translated" anything across all three clips. It kept the drama's English lines, "You got the wrong person" and "under arrest", as spoken. Its mistakes were names and sound-alikes, such as "gelap" (dark) for "gelak" (laugh).
Reference 3: Gemini. In the August round it had at most one blurred phrase on the first talk show, with every party and institution name right. On the TEDx talk it produced 304 coherent words with no hallucinations.
Evidence against our choice
- Soniox Malay-ifies too. In the drama it wrote the English line "You are under arrest" as the Malay "awak ditangkap". The meaning is right and nothing is flipped, but it's the same kind of behaviour we penalized AssemblyAI for, and the original report didn't flag it. The review recommends spot-checking code-switched passages early on in production.
- Doubao isn't a clear second. On the talk shows, its error rate trades places with FunASR and Qwen. One reason we picked it was that FunASR dropped large chunks and Qwen looped, while Doubao did neither. The second half of that turned out to be pinned on the wrong engine (see the last point).
- AssemblyAI is actually very good on the drama: 17.4%, level with Soniox at the top. It only breaks down on dense code-switching in the interviews. It's now a backup.
- Qwen was convicted on two pieces of evidence that both fell apart: the loop was Melia-1's, and the "invented phrase" was the speaker genuinely starting a sentence and correcting themselves. It's now a usable backup but not a standing track, because its timestamps are an evenly spaced grid rather than real timings.
Still open
- Three clips is a small sample.
- No dedicated test of heavy noise with people talking over each other.
- Dense code-switching has only been tested in one setting: political talk shows.
Written by Zeyu Si