<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <title>The Transcribe Labs Blog</title>
  <subtitle>How we test transcription quality, and what we found.</subtitle>
  <link href="https://transcribe.solutions/blog"/>
  <link rel="self" href="https://transcribe.solutions/blog/feed.xml"/>
  <id>https://transcribe.solutions/blog</id>
  <updated>2026-09-22T00:00:00Z</updated>
  <entry>
    <title>Why one engine isn't enough: more engines only help if they fail differently</title>
    <link href="https://transcribe.solutions/blog/engines-that-fail-differently"/>
    <id>https://transcribe.solutions/blog/engines-that-fail-differently</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>One engine gets names and numbers consistently wrong. More engines only help when they fail in different places. A real interview, including where we lost.</summary>
  </entry>
  <entry>
    <title>The most dangerous transcription errors read perfectly</title>
    <link href="https://transcribe.solutions/blog/fluent-errors"/>
    <id>https://transcribe.solutions/blog/fluent-errors</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>A flipped negation reads perfectly and a majority vote won't catch it. So tools should show where they're unsure. We checked 15 of them.</summary>
  </entry>
  <entry>
    <title>WER lies. Here's how we measure accuracy on real interviews</title>
    <link href="https://transcribe.solutions/blog/wer-lies"/>
    <id>https://transcribe.solutions/blog/wer-lies</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>The industry's favourite error metric misleads on interview audio. What we use instead: check the reference first, then count errors that change meaning.</summary>
  </entry>
  <entry>
    <title>When interviews switch languages</title>
    <link href="https://transcribe.solutions/blog/mixed-language-interviews"/>
    <id>https://transcribe.solutions/blog/mixed-language-interviews</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Spanish interviewee, Chinese researcher, an interpreter in between. Vendor docs and lab tests misled us; a real recording gave the opposite answer.</summary>
  </entry>
  <entry>
    <title>Who said that? Speaker attribution when many people talk</title>
    <link href="https://transcribe.solutions/blog/who-said-that"/>
    <id>https://transcribe.solutions/blog/who-said-that</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Built-in speaker labels drift as people are added. Learning each voice first got us above 99% on public meeting data, as long as people take turns.</summary>
  </entry>
  <entry>
    <title>Field notes from 27 languages: how we picked the engines for each</title>
    <link href="https://transcribe.solutions/blog/27-languages"/>
    <id>https://transcribe.solutions/blog/27-languages</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>No engine is best everywhere. The index to our language-by-language tests: which engines we picked for each of 27 languages, and why.</summary>
  </entry>
  <entry>
    <title>Arabic: people speak dialect, subtitles write MSA. How we picked engines anyway</title>
    <link href="https://transcribe.solutions/blog/languages-ar"/>
    <id>https://transcribe.solutions/blog/languages-ar</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Same engine, different reference transcript: error rate went from 49% to 18%. Arabic diglossia forced us to judge engines by meaning, sentence by sentence.</summary>
  </entry>
  <entry>
    <title>Czech: the two-letter prefix that decided our engine lineup</title>
    <link href="https://transcribe.solutions/blog/languages-cs"/>
    <id>https://transcribe.solutions/blog/languages-cs</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>In Czech, &quot;not&quot; is often just ne- on the front of a verb. In our hardest clip, Speechmatics lost it and Soniox kept it. How we picked Czech engines.</summary>
  </entry>
  <entry>
    <title>Danish: we swapped out the cleanest engine on the most interview-like clip</title>
    <link href="https://transcribe.solutions/blog/languages-da"/>
    <id>https://transcribe.solutions/blog/languages-da</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>On the Danish clip closest to a real interview, AssemblyAI was the cleanest engine. We replaced it with Soniox anyway. Our reasoning, and the case against it.</summary>
  </entry>
  <entry>
    <title>German: when Viennese dialect made the steadiest engine switch to English</title>
    <link href="https://transcribe.solutions/blog/languages-de"/>
    <id>https://transcribe.solutions/blog/languages-de</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Three German films and one Austrian. AssemblyAI was cleanest on standard German, then switched to English on Viennese dialect. Doubao never left German.</summary>
  </entry>
  <entry>
    <title>Greek: lose the accent marks and half the words fall apart</title>
    <link href="https://transcribe.solutions/blog/languages-el"/>
    <id>https://transcribe.solutions/blog/languages-el</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>One engine produced real Greek but swallowed every accented vowel whole. Greek error rates run high by nature. How we picked engines on four film clips.</summary>
  </entry>
  <entry>
    <title>English: two reference tracks that fail in different places</title>
    <link href="https://transcribe.solutions/blog/languages-en"/>
    <id>https://transcribe.solutions/blog/languages-en</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four English film scenes read line by line. AssemblyAI and Doubao rarely make the same mistake, which is why both back up ElevenLabs.</summary>
  </entry>
  <entry>
    <title>Spanish: the engine our own code made look broken</title>
    <link href="https://transcribe.solutions/blog/languages-es"/>
    <id>https://transcribe.solutions/blog/languages-es</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four film scenes from Spain, Mexico and Argentina. Doubao was once written off as broken in Spanish. The real cause was a language code we got wrong.</summary>
  </entry>
  <entry>
    <title>Finnish: near-perfect on clear speech, invented lines under gunfire</title>
    <link href="https://transcribe.solutions/blog/languages-fi"/>
    <id>https://transcribe.solutions/blog/languages-fi</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>On clear Finnish conversation, ElevenLabs made no real errors. Under war-film gunfire, every engine invented lines. And why Finnish is where we skip Soniox.</summary>
  </entry>
  <entry>
    <title>French: three engines, one sentence, the same wrong meaning</title>
    <link href="https://transcribe.solutions/blog/languages-fr"/>
    <id>https://transcribe.solutions/blog/languages-fr</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four film scenes from the Paris banlieue, Paris and Quebec. Doubao was the only non-main engine that never invented or flipped a line, and WER told us nothing.</summary>
  </entry>
  <entry>
    <title>Hungarian: why we kept the engine that wrote &quot;ticket&quot; 23 times</title>
    <link href="https://transcribe.solutions/blog/languages-hu"/>
    <id>https://transcribe.solutions/blog/languages-hu</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Five Hungarian film clips. AssemblyAI looped one word 23 times in a brawl, then stayed clean at a snack bar. We kept it. Here is the evidence against that call.</summary>
  </entry>
  <entry>
    <title>Indonesian: someone screams for help, the engine writes 'thank you'</title>
    <link href="https://transcribe.solutions/blog/languages-id"/>
    <id>https://transcribe.solutions/blog/languages-id</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>All four Indonesian test clips come from one film. What engines invented in the noisy scenes, and why Soniox replaced AssemblyAI as a reference.</summary>
  </entry>
  <entry>
    <title>Italian: when our main track heard dismay as praise</title>
    <link href="https://transcribe.solutions/blog/languages-it"/>
    <id>https://transcribe.solutions/blog/languages-it</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four film scenes from Rome, Sicily, Naples and Milan. On one key line, Doubao was the only engine of ten to get it right; ElevenLabs flipped the meaning.</summary>
  </entry>
  <entry>
    <title>Japanese: the engine that looked best was deleting the fillers</title>
    <link href="https://transcribe.solutions/blog/languages-ja"/>
    <id>https://transcribe.solutions/blog/languages-ja</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Two 1950s films and two modern interviews: how filler words flipped our Japanese ranking, why Soniox replaced AssemblyAI, and what is still unresolved.</summary>
  </entry>
  <entry>
    <title>Korean: three engines, one identical wrong negative</title>
    <link href="https://transcribe.solutions/blog/languages-ko"/>
    <id>https://transcribe.solutions/blog/languages-ko</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four Korean film scenes: why ElevenLabs led by a wide margin, why Doubao and Soniox replaced AssemblyAI, and what a three-way shared error taught us.</summary>
  </entry>
  <entry>
    <title>Malay: the speaker switched to English, and the engine translated it back</title>
    <link href="https://transcribe.solutions/blog/languages-ms"/>
    <id>https://transcribe.solutions/blog/languages-ms</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Malay interviews switch between Malay and English mid-sentence. The riskiest error: an engine quietly rewriting the English as fluent Malay.</summary>
  </entry>
  <entry>
    <title>Dutch: no engine crossed a red line. So how do you choose?</title>
    <link href="https://transcribe.solutions/blog/languages-nl"/>
    <id>https://transcribe.solutions/blog/languages-nl</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four Dutch clips, four engines, no invented sentences, flips or loops. When red lines can't separate engines, we look at whose mistakes differ.</summary>
  </entry>
  <entry>
    <title>Norwegian: &quot;we weren't unfaithful&quot; became &quot;we chose to be unfaithful&quot;</title>
    <link href="https://transcribe.solutions/blog/languages-no"/>
    <id>https://transcribe.solutions/blog/languages-no</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Five Norwegian film clips and a podcast. One engine turned &quot;we weren't unfaithful&quot; into &quot;we chose to be unfaithful&quot; three times, on clear audio. How we chose.</summary>
  </entry>
  <entry>
    <title>Polish: drop one &quot;nie&quot; and &quot;I won't have kids&quot; becomes &quot;I will&quot;</title>
    <link href="https://transcribe.solutions/blog/languages-pl"/>
    <id>https://transcribe.solutions/blog/languages-pl</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four Polish film clips, one flipped negation, and that was enough to keep an engine out. How we picked engines for Polish, and why we swapped one.</summary>
  </entry>
  <entry>
    <title>Portuguese: three flipped lines, and a key sentence rescued</title>
    <link href="https://transcribe.solutions/blog/languages-pt"/>
    <id>https://transcribe.solutions/blog/languages-pt</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Brazilian film scenes plus three real conversations: Gemini flipped a line in each, yet often beat the main track. European Portuguese has no test clip.</summary>
  </entry>
  <entry>
    <title>Romanian: &quot;Are you hungry?&quot; became &quot;I'm starving&quot;</title>
    <link href="https://transcribe.solutions/blog/languages-ro"/>
    <id>https://transcribe.solutions/blog/languages-ro</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>On the clearest Romanian clip, one engine flipped three meanings in a minute. Romanian was also the first language where results made us swap a track.</summary>
  </entry>
  <entry>
    <title>Russian: the line both reference engines got wrong</title>
    <link href="https://transcribe.solutions/blog/languages-ru"/>
    <id>https://transcribe.solutions/blog/languages-ru</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>How we tested eight speech engines on four hard Russian film scenes, what went wrong, and why we run ElevenLabs with AssemblyAI, Doubao and Gemini.</summary>
  </entry>
  <entry>
    <title>Swedish: when &quot;apologize!&quot; became &quot;I apologize&quot;</title>
    <link href="https://transcribe.solutions/blog/languages-sv"/>
    <id>https://transcribe.solutions/blog/languages-sv</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four Swedish test clips, nine engines. In one scene a bully orders his victim to apologize; four engines turned it into &quot;I apologize.&quot; Only one got it right.</summary>
  </entry>
  <entry>
    <title>Thai: part of our primary engine's score was our own bug</title>
    <link href="https://transcribe.solutions/blog/languages-th"/>
    <id>https://transcribe.solutions/blog/languages-th</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Three Thai clips: a bug that stripped vowel and tone marks, Gladia's runaway repeats, why best-scoring Qwen isn't used, and what we still don't know.</summary>
  </entry>
  <entry>
    <title>Turkish: how one model upgrade broke a clean record</title>
    <link href="https://transcribe.solutions/blog/languages-tr"/>
    <id>https://transcribe.solutions/blog/languages-tr</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four Turkish film scenes: why AssemblyAI's newer model did worse, why last-placed Doubao stays, why Soniox is only a backup, and Gemini's missing endings.</summary>
  </entry>
  <entry>
    <title>Ukrainian: 54 seconds of silence, turned into a lecture</title>
    <link href="https://transcribe.solutions/blog/languages-uk"/>
    <id>https://transcribe.solutions/blog/languages-uk</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>A stretch with only a TV in the background, and two engines wrote dialogue into it. Why we turned down Soniox for Ukrainian, and the risk we kept.</summary>
  </entry>
  <entry>
    <title>Vietnamese: the 'hallucination' was a market vendor</title>
    <link href="https://transcribe.solutions/blog/languages-vi"/>
    <id>https://transcribe.solutions/blog/languages-vi</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>We have only two valid Vietnamese test clips. What they show, why we run ElevenLabs with Doubao, Soniox and Gemini, and what two clips can't prove.</summary>
  </entry>
  <entry>
    <title>Mandarin: choosing engines through gunfire and screaming</title>
    <link href="https://transcribe.solutions/blog/languages-zh"/>
    <id>https://transcribe.solutions/blog/languages-zh</id>
    <published>2026-09-22T00:00:00Z</published>
    <updated>2026-09-22T00:00:00Z</updated>
    <author><name>Zeyu Si</name></author>
    <summary>Four Mandarin film scenes, twelve engines, read line by line: who invents sentences in noise, who gives up, and why Chinese ended up with four reference tracks.</summary>
  </entry>
</feed>
