A voice assistant that works perfectly in a quiet office can fall apart on its first real call: a farmer phoning from a field, a caller who starts in Hindi and finishes in English, a grandmother who speaks Bhojpuri and expects to be understood. Building for India means building for all of them.
This article walks through how we approach multilingual voicebots at Eurys — the problems that show up in production, the pipeline we use to handle them, and the checks we run before any assistant goes live on a public helpline.
Why Indian languages are hard for voice AI
Most speech models are trained on clean, single-language audio. Real helpline traffic looks nothing like that. Three problems come up again and again.
- Telephony audio. Phone calls arrive at 8 kHz, often compressed, with background noise from traffic, markets and crowded homes. Models tuned on studio-quality speech lose a large share of their accuracy here.
- Dialect and accent range. Hindi spoken in Lucknow, Patna and Indore sounds different. Tamil, Telugu and Bengali each carry strong regional variation. A single "standard" model rarely covers the spread.
- Domain vocabulary. Scheme names, district names, application numbers and medical terms are rare in general training data, yet they are exactly what callers say most.
A caller asks about their "PM Kisan ki kist" on a noisy line. A general model hears unrelated words, the bot asks them to repeat three times, and they hang up.
A telephony-tuned speech model with a custom vocabulary of scheme names, plus an intent layer that understands the request even when one word is misheard.
The pipeline, stage by stage
A voicebot is not one model. It is a chain, and each link can lose meaning. We design every stage to pass along not just its best guess, but how confident it is.
- Language identification. The first two to three seconds of speech decide which recognition model to use. We let callers switch language at any point, so identification keeps running throughout the call.
- Speech recognition. Models are fine-tuned on telephony audio for each supported language and boosted with a domain word list for the specific deployment.
- Understanding. A language model maps the transcript to an intent and extracts details like dates, IDs and locations — tolerating spelling variation and partial matches.
- Response and speech. Replies are written in short, spoken-style sentences, then voiced with a natural text-to-speech voice in the caller's language and register.
Design tip: keep spoken prompts under 15 words where possible. Callers cannot scroll back on a phone line, so long menus get forgotten before they end.
Handling code-mixing
Code-mixing — switching languages within a sentence — is the norm, not the exception. "Mera application status check karna hai" mixes Hindi grammar with English nouns, and callers expect the bot to answer in the same blend.
We handle this in three ways. Recognition models are trained on mixed-language transcripts rather than separated corpora. Entity extraction normalises words written in either script, so "application" and "एप्लीकेशन" resolve to the same thing. And response generation mirrors the caller's mix instead of forcing formal, single-language replies that feel stiff.
caller: "Mera certificate download nahi ho raha"
intent: certificate_download_failed
reply: "Koi baat nahi. Main aapko WhatsApp par
certificate ka link bhej deta hoon."
The goal is not a bot that speaks perfect Hindi. It is a bot that speaks the way the caller does. Eurys Engineering
Testing with real callers
Lab accuracy numbers are a starting point, not a launch signal. Before an assistant goes live, we run it against recorded and live pilot traffic and review failures by hand.
What we review
- Calls where the caller repeated themselves more than once
- Calls transferred to a human agent, and the reason for each transfer
- Calls that ended within the first 20 seconds
- A random sample of "successful" calls, to catch confident mistakes
Each failure feeds back into the domain vocabulary, the intent examples or the prompt wording. Most improvements in the first month come from this loop rather than from changing models.
What good looks like
Numbers vary by use case, but these are the signals we look for once a multilingual assistant has settled in.
Just as important is what callers do not notice: they are not asked to pick a language from a long menu, they are not forced to speak formally, and when the bot does need help, a human picks up with the full conversation already on screen.
A launch checklist
If you are planning a multilingual voice assistant, these are the questions to answer before going live.
- Which languages and dialects make up 90% of your expected callers?
- Do you have a domain vocabulary list — scheme names, places, product terms?
- Can callers switch language mid-call without starting over?
- Is there a clear path to a human agent, with context passed along?
- Who reviews failed calls each week, and how do fixes get deployed?
Get these right and the model choice matters far less than you might expect. Want to hear what this sounds like in practice? Request a demo and we will set one up in your language.



