Skip to main content

Part of: 27 languages, one by one

Arabic: people speak dialect, subtitles write MSA. How we picked engines anyway

· Zeyu Si

On this page

One engine, one level of skill. On one Arabic recording its error rate was 49.2%. On another, 18.2%.

The engine didn't change. The reference transcript did. The first one was written in formal Arabic. The second recorded, word for word, the dialect the speaker actually used.

The test clips: two recordings, two dialects

  • Talhouk: a public talk by a Lebanese speaker about mother tongue and Arabic-language education. He speaks colloquial Lebanese Arabic. The subtitles render it in Modern Standard Arabic (MSA).
  • Fenjan (روقي): a current-affairs interview in Gulf (Saudi) Arabic. The subtitles transcribe the spoken dialect verbatim.

That's all we have. Free material that combines dialect audio, human-made subtitles and verbatim fidelity is very hard to find, and we stopped after these two. They were chosen for dialect, not difficulty. Our hardest tier (people talking over each other, heavy background noise) isn't covered for Arabic.

And only the second clip can be used to rank engines. Here's why.

Engines we tested

Ten: ElevenLabs Scribe v2, AssemblyAI, Speechmatics (enhanced and Melia-1), Doubao (ar-SA), Soniox, Deepgram, Gladia, FunASR (fun-asr-mtl) and Qwen (qwen3.5-omni). Gemini was added in a follow-up round in August 2026.

Red lines: Arabic turned out to be a "safe" language

On the Fenjan interview we read eight transcripts line by line: ElevenLabs, Soniox, Speechmatics, Qwen, Doubao, Gladia, Deepgram and FunASR. All of them kept the meaning. No invented sentences, no flipped negations, no fabricated passages.

The problems were elsewhere:

  • AssemblyAI looped on a stretch of the Talhouk talk and produced content that isn't in the audio, attached to timestamps of zero length. It was the only genuine hallucination in the whole Arabic panel.
  • FunASR slipped an English word, "Also", into an Arabic sentence (a small language drift) and turned the year 1990 into 2990.
  • Doubao and FunASR scored 41.4% and 37.7% on Fenjan, roughly double the leaders. Reading closely, most of those "errors" came from two things: the hamza mark left off, and years spelled out the way people say them in dialect (2011 as الفين واحدعش). A reader wouldn't be misled. We first concluded Doubao was weak in Arabic; we later withdrew that.

On the same clip, Speechmatics scored 15.7%, Qwen 17.9%, Soniox 18.2% and ElevenLabs 20.1%. A few points apart.

What we chose, and why

Primary: ElevenLabs Scribe v2. On Fenjan it sits at the back of the leading group, so Arabic isn't its strongest language. But it was faithful on both clips. On the Lebanese talk it kept dialect words like بدي, هيدا and وين almost untouched. On the Gulf interview it wrote the more generic أيش where the speaker said وش ("what"), without losing the meaning. It's our primary engine in every language mainly for speaker separation and stability on long recordings, not for this particular score.

Reference 1: Soniox. No hallucinations on either clip, and the most faithful reading of the Gulf dialect. On the Lebanese talk it did the opposite: it wrote the dialect as formal Arabic, which is why its WER there was the lowest. So it errs in the opposite direction from our primary engine. One leans toward dialect, the other toward the written form. That's more useful for cross-checking than two engines that go wrong the same way.

Reference 2: Speechmatics. Lowest error rate on Fenjan, clean output, no breakdowns. It's a conventional acoustic model, built differently from Soniox. Its transcripts come without punctuation, which is an API setting and doesn't affect the words.

Reference 3: Gemini. In the August round, Arabic went into the "can be added" group: no red lines on at least two conversational or speech clips.

Not chosen: Qwen was as accurate as the top two, but it's a large-language-model transcriber like Gemini and doesn't give word-level timestamps, so it stays on the bench. Doubao works differently from everything else in the lineup; it's a possible backup but isn't wired in.

Evidence against our choice

  • We once listed "Soniox is the most faithful to Lebanese dialect" as a reason for picking it. A word-by-word re-check showed we had it backwards. The engines that kept the dialect words were ElevenLabs, FunASR, Qwen and Doubao. Soniox was the one that formalized most. That reason is gone; Soniox stays for the reasons above.
  • The only trustworthy ranking rests on a single recording. Speechmatics, Qwen and Soniox are two or three points apart, and one clip can't separate them.
  • Our first pass on Arabic relied mainly on WER and output structure. Reading for meaning came later, and it overturned the "Doubao and FunASR are weak in Arabic" verdict.

Still open

  • Only Levantine and Gulf Arabic were tested. Egyptian and other dialects aren't covered.
  • Once we find a third and fourth dialect-faithful recording, the ranking needs to be rerun.
  • Both clips are a public talk and a media interview, not research interviews.

Our Arabic transcription page

Written by Zeyu Si