Back to Blog
    voice-to-text
    speech-to-text
    voice-dictation
    voice-transcription
    audio-transcription

    Best Bilingual AI Dictation App: A Russian-English Stress Test

    Burlingame, CA
    Best Bilingual AI Dictation App: A Russian-English Stress Test

    The hardest bilingual dictation is rarely a clean sentence in one language followed by a clean sentence in another. Real people switch halfway through a thought. They borrow a noun, bend it around another language's grammar, then switch back without warning.

    We tested AI Dictation, Wispr Flow, Superwhisper, and ten speech recognition models on a 70-second Russian-English clip that does exactly that. GPT Transcribe produced the strongest practical result for preserving both languages, and we now use it in AI Dictation's cloud mode.

    Disclosure: AI Dictation is our product. The competing app outputs are reproduced as captured, the scoring method is described below, and the full model table includes results that beat us on individual metrics.

    A microphone sending multilingual speech through three transcription systems

    The test audio: "Смотря какой fabric"

    The source is a short YouTube clip titled "Смотря какой fabric смотря сколько details". A Russian-speaking woman in New York discusses fashion in what many Russian speakers would call Rushlish. Russian grammar carries the sentence, but English words keep appearing inside it: "Mango colors," "capri pants," "look couture," "fabric," "price," and "spring and summer."

    That makes the clip a nasty speech-to-text test. A system has to recognize both languages, keep the English in Latin script, avoid translating the Russian, preserve the speaker's wording, and reach the final sentence. The designer's name adds another trap. The correct name is Simon Chang, but it is inflected inside a Russian sentence.

    The human-annotated baseline

    We created a 146-word reference transcript by listening to the recording and preserving the code-switching as spoken. This is a transcription, not a translation.

    Что модно в этом сезоне, на ваш взгляд и на взгляд Simon Чена? Mango colors, шелк очень модный. Очень модные пиджаки с какими-то details, брюки модно, capri pants модно. И это всё on design. Джинсы он очень красивые делает. Можно одеть capri джинсы, красивую блузку под низ, накинуть jacket. Можно так на работу выйти, можно так пойти и погулять потом. Так что он universal. И модно всё. Всё красивое и fit его amazing, amazing на всех. У него look couture без этих couture prices. Его костюмы, смотря какой fabric, смотря откуда приходит fabric, смотря сколько details в этом пиджаке, может price от 225, самый дорогой 380. Он very очень-очень affordable. Ну что ж, замечательно. Ну и заключение, чисто традиционный вопрос: ваше пожелание нашим милым и обаятельным телезрительницам? Я вам желаю хорошей spring and summer, чтобы вы приходили красиво одевались, и мы вам можем помочь.

    We scored two numerical measures. Normalized word error rate, or WER, counts substitutions, deletions, and insertions against the reference. Lower is better. English-token recall checks how many of the 27 English words embedded in the Russian speech survived in Latin script. Higher is better.

    Numbers alone miss some ugly failures, so we also checked completeness, proper names, punctuation, transliteration, and translation.

    App results at a glance

    App or pipelineWEREnglish-token recallComplete recordingMain failure
    AI Dictation with GPT Transcribe13.7%96.3%YesMissed Simon Chang
    Wispr Flow27.4%63.0%YesTransliterated English words and damaged the closing sentence
    Superwhisper S1 Voice, raw output46.6%44.4%NoOmitted a large middle section
    Superwhisper S1 Voice plus S1 LanguageNot scoredNot retainedYesNormalized the mixed speech toward Russian

    The app ranking is straightforward for this recording. AI Dictation preserved the most bilingual content. Wispr Flow finished the clip but converted several English words into Russian-looking phonetics. Superwhisper's raw voice model retained some code-switching but lost too much audio, while its standard language stage removed the very behavior we wanted to measure.

    Wispr Flow preserved the shape, then lost the words

    Wispr Flow returned a complete transcript and kept visible English phrases such as "Mango colors," "details," "look couture," and "spring and summer." That is better than a pipeline that translates everything into one language.

    The trouble appears in the words that sit near the boundary between Russian and English. "capri pants" became "Капри-пенс." "fit" became "фет." "fabric" became "фабрик." The closing phrase "мы вам можем помочь" became "мы вам можем получить," changing "we can help you" into a sentence that does not make sense in context. Simon Chang also became "Саймон Чены."

    Wispr Flow result for the Russian-English Rushlish benchmark

    Those errors explain the 63.0% English-token recall. Wispr Flow heard much of the content, but it treated several English words as sounds to spell in Cyrillic. For a bilingual user, that means manual repair after nearly every switch.

    Wispr Flow still has a real advantage for people who want a polished cross-device dictation product. Our result only addresses one specific job: preserving dense Russian-English code-switching. Readers comparing the broader category can see our Wispr Flow alternatives guide.

    Superwhisper depended heavily on its two-model pipeline

    We tested Superwhisper with its "Super" preset, automatic language selection, S1-Voice, and S1-Language. That is the normal configured experience shown in the app, not a hand-picked external model.

    Superwhisper Super preset using S1-Voice and S1-Language

    The raw S1-Voice output caught a few mixed phrases, including "Simon Chen," "amazing," "look couture," and "spring and summer." It also dropped most of the first half of the recording. The output contained 95 words against the 146-word human reference, producing 46.6% WER and 44.4% English-token recall.

    The S1-Language stage made the prose more internally consistent, but it did so by translating or normalizing the mixed speech toward Russian. That might be useful when the desired output is monolingual Russian. It fails a fidelity test where the English words are part of the speaker's voice.

    This distinction matters. Superwhisper exposes both a voice model and a language model because the final text is the result of both. Judging S1-Voice alone would hide what users receive after the standard pipeline runs. Our older AI Dictation vs Superwhisper comparison discusses the rest of their product differences.

    AI Dictation with English and Russian selected

    AI Dictation lets users select more than one transcription language. For this run, English and Russian were both enabled.

    AI Dictation language settings with English and Russian selected

    The GPT Transcribe result preserved 26 of the reference's 27 English tokens. It kept "Mango colors," "capri pants," "on design," "capri jeans," "jacket," "fit," "look couture," both instances of "fabric," the prices, and the final "spring and summer."

    Here is the complete output:

    Что модно в этом сезоне, на ваш взгляд и на взгляд Саймона Чиане? Mango colors. Шелк очень модно, очень модные пиджаки с какими-то details, брюки модно, capri pants модно. И это все on design, джинсы он очень красиво делает. Можно одеть capri jeans, красивую блузку под низ, накинуть jacket, и можно так на работу тоже выйти, и можно так пойти погулять потом. Так что он universal, и модно все. Все красивое, и fit его amazing, amazing на всех. Его look couture без этих couture prices. Его костюмы, смотря какой fabric, смотря откуда приходит fabric, смотря сколько details в этом пиджаке, может price от 225, самый дорогой 380. Так оно very, очень-очень affordable. Ну что ж, замечательно. Ну и в заключение чисто традиционный вопрос: ваши пожелания нашим милым и обаятельным телезрительницам. Я вам желаю хорошей spring and summer, чтобы приходили, красиво одевались, и мы вам можем помочь.

    It is not perfect. Simon Chang became "Саймона Чиане," and there are small Russian grammar differences. Still, the sentence remains bilingual, readable, complete, and faithful to the source. That is why the 13.7% WER needs to be read beside the 96.3% English-token recall.

    We tested the models separately too

    An app comparison can tell you which app worked. It cannot tell you why. We ran the same audio through ten model configurations using one generic prompt:

    Transcribe exactly as spoken. Preserve language switching and keep every word in its spoken language. Do not translate, paraphrase, correct, add, or omit words. Output only the transcript.

    We intentionally excluded a Russian-English-specific prompt from the production ranking. That prompt improved one GPT Transcribe run to 12.3% WER, 100% English-token recall, and the correct name Simon Chang. It also gave the model advance knowledge that a general dictation app does not have. Nice lab score, bad product test.

    ModelWEREnglish-token recallSimon Chang correct?
    OpenAI GPT 4o Mini Transcribe12.3%88.9%No
    OpenAI GPT Transcribe13.7%96.3%No
    Gemini 3.1 Pro Preview, low thinking15.1%92.6%Yes
    Gemini 3.5 Flash Lite, minimal thinking15.1%88.9%No
    Gemini 3.6 Flash, medium thinking16.4%88.9%No
    OpenAI GPT 4o Transcribe17.1%63.0%No
    Gemini 3.7 Flash, medium thinking17.8%88.9%No
    Gemini 3.1 Pro Preview, high thinking19.9%92.6%No
    Groq Whisper Large V322.6%70.4%No
    Groq Whisper Large V3 Turbo93.2%74.1%No

    GPT 4o Mini Transcribe won on raw WER by 1.4 percentage points. GPT Transcribe won the metric that mattered more for this product decision. It retained 96.3% of the English words while keeping the Russian intact.

    Gemini 3.1 Pro at low thinking deserves a mention. It was the only generic-prompt run that wrote Simon Chang correctly, and its English recall reached 92.6%. Raising the thinking setting made the result worse, which is a useful reminder that more inference does not automatically improve verbatim transcription.

    Groq Whisper Large V3 produced readable mixed text but translated one full phrase into English and transliterated several English words into Cyrillic. The Turbo variant collapsed badly. It translated most of the Russian into English and even copied part of the prompt into the transcript. Fast is irrelevant when the output is wrong.

    Other provider runs failed in less interesting ways. Together's Parakeet behaved like an English-focused model on this clip. Together Whisper misidentified the Russian. A fal-hosted model translated most of the Russian to English. We could not complete the Replicate run because the test account had no credits, and AWS Transcribe was blocked by the available account's S3 permissions. We report those as incomplete or failed tests, not scores.

    Why GPT Transcribe now powers AI Dictation

    The lowest WER did not win automatically. A bilingual dictation app has a different job from a meeting summarizer or a monolingual transcription service. It must preserve the words people chose, including the language they chose for each word.

    GPT Transcribe gave us the best balance in this recording. It completed the full tail, retained almost every English token, kept the mixed scripts, and responded well to a generic language-preservation prompt. Release v0.0.113 moves AI Dictation's cloud transcription to GPT Transcribe across macOS, iOS, Android, and Windows. The macOS realtime path uses the matching live transcription backend and commits the recording only after the final audio buffer drains.

    That last detail came from a painful bug. During development, committing the live audio every 0.8 seconds turned one sentence into a chain of tiny recognition jobs. Punctuation broke, capitalization wandered, and the final words became unstable. One recording should be one linguistic context. We fixed the commit path before shipping the release.

    If your work is mainly monolingual English, this test should not decide your app choice. Read our broader best voice-to-text software guide. If you dictate Russian, our Russian speech-to-text app comparison covers the normal monolingual case.

    For dense bilingual speech, the result is firmer. AI Dictation with GPT Transcribe was the best bilingual voice-to-text app in this Russian-English benchmark. Wispr Flow remained usable but needed repairs at language boundaries. Superwhisper's standard language pass optimized for consistency and erased too much of the code-switching.

    Want to try the same kind of sentence yourself? Download AI Dictation, select both languages, and dictate the way you actually speak.

    Frequently Asked Questions

    What is the best bilingual AI dictation app?

    AI Dictation gave the strongest Russian-English result in this test after its move to GPT Transcribe. It preserved 96.3% of the English tokens inside Russian speech and completed the full recording.

    Can voice-to-text apps transcribe two languages in one sentence?

    Yes, but accuracy varies sharply. GPT Transcribe preserved Russian-English code-switching well, while some Whisper pipelines translated English words, transliterated them into Cyrillic, or translated most of the recording into English.

    Why did AI Dictation choose GPT Transcribe?

    GPT Transcribe did not post the lowest raw word error rate, but it retained 96.3% of the English tokens, preserved the complete recording, and kept mixed-language phrasing intact more reliably than the other practical app candidates.

    How was this multilingual dictation test scored?

    We compared every result with a 146-word human transcript. We measured normalized word error rate, recall of 27 English tokens spoken inside Russian sentences, completeness, script preservation, and unwanted translation.

    Ready to try AI Dictation?

    Experience fast voice-to-text on your device. Free to download.

    Download Free