Part of: 27 languages, one by one
Japanese: the engine that looked best was deleting the fillers
· Zeyu Si
On this page
On two modern Japanese interview clips, AssemblyAI first looked like the most accurate of the four engines we tried, with a character error rate as low as 8.4% on one of them. Then we worked out where the lead came from. The subtitles we score against drop spoken fillers and hesitation sounds. AssemblyAI drops them too. Soniox writes them down, and every one it wrote counted as an error. Strip the fillers from both sides and Soniox wins all four clips, reaching 2.0% on the last one, lower than our primary engine.
That finding changed who sits on the Japanese reference tracks.
The test clips
Four clips in two pairs:
- Seven Samurai (1954): Tohoku dialect, dense shouting, old-film hiss.
- Rashomon (1950): standard Japanese, but heavy rain and three people interrupting one another.
- A two-person interview from YouTube: frequent turn-taking, clean studio sound.
- An episode of a Japanese-language YouTube show: two hosts constantly backchannelling (そうだね, うん and the like), with speaker changes every few words.
On our three-level difficulty scale the films are top level and the modern clips are middle level. We didn't test a separate easy Japanese clip (one or two people, steady pace, no noise).
For the modern clips we scored against the uploader's hand-made subtitles, never auto-captions. The first interview we picked turned out, on a line-by-line check, to be spoken English with Japanese translation subtitles. Scoring a transcript against a translation would have been meaningless, so we swapped it out.
Engines we tried
ElevenLabs Scribe v2, AssemblyAI, Doubao and Speechmatics in the first round. Soniox, Gladia, Deepgram, FunASR and Qwen in follow-up runs. Gemini was tested separately.
Where the engines crossed red lines
Repetition loops. In the Seven Samurai shouting, AssemblyAI repeated one sound 95 times. On the same clip Gladia repeated 447 times. Once a transcript fills up with one syllable, whatever came after is gone.
Invented sentences. Also on Seven Samurai, AssemblyAI made up whole sentences, and on Rashomon it tacked unspoken words onto the end of a line.
Inserted content. On the old films Doubao once wrote out a violent command that nobody gave. Its clean record only holds for the modern clips.
Flipped meaning. Soniox reversed the sense of two lines in the dialect-and-shouting stretch of Seven Samurai. On that same clip AssemblyAI was looping and inventing sentences. Dropping Soniox for those two errors while keeping AssemblyAI would be judging them by different rules, so we treated it as a problem of that scene.
Language drift. When we tried Gemini as a stand-in primary engine, its Japanese came out with some kanji in simplified Chinese forms: 绍介, 连络, 结婚 instead of 紹介, 連絡, 結婚. Right characters, wrong script.
Substitutions (not red lines, but misleading): Speechmatics heard 人前式, a wedding held before family and friends, as 人権意識, "human rights awareness". Doubao and AssemblyAI both wrote 酌み交わす ("share a drink") as its homophone 組み交わす.
What we chose
Primary: ElevenLabs Scribe v2. The only engine that never collapsed on any of the four clips, and 5.1% on the YouTube show.
References, in order: Soniox, Doubao, Gemini. All three run on every Japanese file.
- Soniox: wins all four clips once fillers are stripped. We put its two reversals on the old film down to the scene.
- Doubao: we first thought it couldn't do Japanese because it returned garbage. We simply hadn't sent the right language code. With ja-JP it was clean on the modern clips, third on Rashomon, and mid-table but intact on Seven Samurai. It comes from a very different lineage than Soniox, so the two rarely fail in the same spot.
- Gemini: no red-line errors on the two modern clips.
AssemblyAI became a backup.
Evidence against our choice
- AssemblyAI really is accurate on modern interviews, and on clean, real conversation its problems mostly vanish. We replaced it mainly because of the looping and invented sentences on old films, and because it's no longer first once fillers are stripped.
- An earlier round reached the opposite conclusion. At that point Soniox didn't look better than the two incumbents on interview material, and the rule was "no change on a tie". The filler discovery reversed it. So the decision rests on one scoring choice: whether fillers count.
- Gemini made a lot of mistakes on the two old films. In Seven Samurai it turned 米 ("rice") into the surname 亀岡. In Rashomon it wrote the Heian-era police title 検非違使 as 蛇石 ("snake stone"), turned "enough preaching" into "enough coughing", and dropped the final sentence. Rashomon is clear dialogue, so the trouble tracks recording quality more than noise. If a client's recording is poor, this track may slip with it.
Open problems
- We have only two real interviews, and both are clean. We have no scored example of a modern Japanese interview in a noisy room.
- If a client wants strict verbatim with every えーと kept, Soniox's habit of writing fillers is a strength. If they want a cleaned-up transcript, the basis for the ranking needs another look.
- The simplified-kanji problem showed up when Gemini stood in for the primary engine. We have no separate data on whether it happens when Gemini runs as a reference track.
Our Japanese transcription page
Written by Zeyu Si