Hinglish Is Not a Language Setting: Code-Mixed Speech, Accuracy and AI Voice Agent Latency
Vicky Yadav
Backend & AI Systems Engineer, Zoice

A caller says, “Mera EMI ka due date change karna hai, can you check?” The sentence switches between Hindi and English twice, and many listeners in India wouldn't find that unusual. Many voice agents struggle with it: key words come back wrong, the language model answers a question nobody asked, the agent asks for a repeat, and one quick exchange becomes several slow ones. That is why code-mixed speech is as much an AI voice agent latency problem as an accuracy problem.
The usual first fix is a dropdown: set the agent's language to Hindi or to English and hope for the best. Hindi-English code-mixed speech, usually called Hinglish, fits neither option, and it is not simply a third option in a conventional language selector either. It is a way of speaking in which the speaker picks, phrase by phrase, whichever language says the thing best. This post walks through what code-mixing does to each stage of a voice pipeline, from recognition to the turn timing callers notice, and what to design for instead.
Why Code-Mixed Speech Breaks Speech Recognition
The organisers of the MUCS 2021 speech recognition challenge describe code-switching as speech in which “multiple languages are freely interchanged within a single sentence or between sentences.” In practice it is messier. A switch can happen mid-phrase (“payment ho gaya but receipt nahi aaya”). English nouns take Hindi grammar (“files ko upload karo”). Everyday borrowed words such as “account”, “EMI” or “recharge” are pronounced with sounds that fall somewhere between the two languages.
A recogniser set to one language has strong expectations about which words are likely to come next. Set it to Hindi and an English word like “statement” may be bent into a Hindi-sounding word or spelt out phonetically in Devanagari. Set it to English and the small Hindi words that hold the sentence together, such as “nahi”, “kab” and “kyun”, get mapped onto English words that sound vaguely similar. In practice these errors often hit the words that matter most to the agent: the negation, the question word, the product name.
Architecture matters too. A conventional multilingual pipeline may detect the language first, route the audio, and send it to a language-specific recogniser, which has to commit to one language even when the caller is using two. A code-switching-aware recogniser instead transcribes mixed-language speech directly, without a separate routing step. Not every multilingual system works by routing, so ask a vendor which design you are getting.
The problem is not new. MUCS 2021 covered code-switching in two pairs, Hindi-English and Bengali-English. Its end-to-end baseline reported a 32.45% average word error rate across the two code-switching test sets: 27.7% for Hindi-English and 37.2% for Bengali-English. Word error rate (WER) is roughly the share of words transcribed wrongly, so the average is close to one word in three.
Models have improved in the five years since, but recent evaluations still show a very wide spread. SwitchLingua, a code-switching dataset and benchmark presented at NeurIPS 2025, covers 12 languages and more than 80 hours of recorded audio. On its Hindi-English test set, the reported WER ranged from 0.2697 for Whisper-Large-v3 to 0.6953 for Qwen2-Audio-7B-Instruct on the same audio.
Read those figures with care. They come from the paper's results table and from offline (batch) evaluation, not from live streaming calls, and they describe the models the authors tested rather than every Hinglish system in production today. What they do show is how much the choice of model matters, which makes it the first decision to get right: a model that does well on clean Hindi or clean English tells you little about a caller who uses both in one breath.
Key Insight
Code-mixed speech costs latency twice: once in processing, and again in repair turns, when the agent mishears and has to ask again. Choose models that handle code-mixing natively, decide your transcript script on purpose, and measure whole turns, repairs included, on your own calls.
On this page
Transcript Script Is a Design Decision
Even a model that hears every word correctly still has to decide how to write it down, and for Hinglish there is no single right answer. As Deepgram's guide to Hinglish recognition puts it, there is no standard way to write Hinglish: the same Hindi words show up in Devanagari, in Roman letters or in a mixture, depending on who is typing and where. “Payment kab hoga” and “पेमेंट कब होगा” are the same utterance, and so is any mix of the two.
Some speech-to-text providers now offer this as an explicit choice. Sarvam's documentation describes five output modes for its speech-to-text, and three of them matter for code-mixed audio: “transcribe” writes everything in native script, “codemix” keeps English words in English, and “translit” outputs Roman script. The point is that script is a decision you make for the systems that read the transcript next; a language setting does not make it for you.
Three consumers of the transcript usually matter:
- The language model. Test which form your chosen model reasons over most reliably; mixed script that matches how the words were spoken often keeps English product terms intact.
- Structured extraction and your CRM. An order ID, a plan name or an email address must come out in Roman script whatever language the caller used, so normalise these fields explicitly.
- Retrieval. If retrieval relies heavily on lexical or script-sensitive matching, a query written in a different script from the indexed content can quietly reduce recall. Consistent script conventions, language metadata, multilingual embeddings or query normalisation can all help, so test retrieval against the transcript convention you actually use.
Script also distorts measurement. Word error rate compares strings, so a transcript that writes “payment” in Devanagari counts as wrong against a Roman-script reference, even though the word was heard correctly. The SwitchLingua authors propose a new metric, Semantic-Aware Error Rate (SAER), which “combines the semantic error and language-specific form error into a weighted equation”. They do so because WER-style scoring captures code-switching errors poorly. If you are comparing vendors, normalise script before comparing WER, and never put figures from different test sets side by side.
Core Components
A speech-to-text model that transcribes Hindi and English in the same utterance, chosen by testing on your own recorded calls rather than monolingual benchmarks.
One deliberate rule for writing Hinglish down, with identifiers such as order IDs and plan names always normalised to Roman script.
A language model and prompt tested on real code-mixed transcripts and told to mirror the caller's mix of languages.
Language-tagged knowledge-base chunks, with queries normalised to the index's script or matched through multilingual embeddings, tested on real Hinglish questions.
End-to-end turn latency, from the caller stopping to the agent's voice starting, tracked at the 95th percentile alongside repair turns.
Where Code-Mixing Shows Up in AI Voice Agent Latency
Voice-agent latency targets are usually set against human conversation. In a study of turn-taking across ten languages (none of them Indian), Stivers and colleagues timed answers to yes/no questions in everyday conversation and found a mean response offset of about 208 ms across the full dataset: the gap between the question ending and the answer starting. The language-specific means fell within roughly ±250 ms of that overall mean. This is a human baseline, not a specification for machines, but it explains why callers notice delays that look small on a dashboard.
Streaming recognisers have improved, and code-mixing's delay often hides outside the speech-to-text model itself. Sarvam says the Fast mode of its streaming Saaras V3 recogniser guarantees under 150 ms time to first token (TTFT). That is the time until the first piece of transcript arrives, not the full time from the caller stopping to the agent's voice starting. Saaras V4, which Sarvam released in September 2026 as its latest generation, also streams with a TTFT below 150 ms and is described as robust to noisy audio and code-mixing. Saaras V3 remains available, and Sarvam's documentation still points to it for 8 kHz telephony audio. All of these are vendor-reported figures.
Accuracy gains don't always carry over to streaming either. In figures reported by Deepgram rather than an independent benchmark, an update to Nova-3 Multilingual cut mean WER by about 34% (relative) in batch mode but only about 21% in streaming, across its supported languages.
In our engineering view, the cost that is easiest to miss is the repair turn. When a misheard negation or product name sends the language model the wrong way, the agent either answers the wrong question or asks the caller to repeat. Illustratively, with round numbers made up for this example rather than measured: a recogniser that is 100 ms faster saves under a second across an eight-turn call, while one misheard “nahi” that forces the caller to repeat themselves can cost five seconds or more in a single exchange. The table below sketches where code-mixing tends to cost time and correctness at each stage; it is a qualitative guide, not measured data, and the real effect depends on your models, your audio quality and how your callers speak.
| Stage | What code-mixing does | What to watch |
|---|---|---|
| Language detection and endpointing | A detector that waits for enough audio to decide “Hindi or English” can hold the turn, then reverse its decision when the next word switches | Avoid designs that must commit to one language before transcribing |
| Speech-to-text | Words from the other language get forced into the configured vocabulary, and streaming partial transcripts may be rewritten more often | WER on your own recorded calls, scored after script normalisation |
| Language model | Garbled or inconsistent-script input weakens reasoning and invites wrong answers | Answer correctness on real code-mixed transcripts, not on translated test sets |
| Retrieval | If retrieval leans on lexical or script-sensitive matching, a query in a different script from the index can miss the right passage | Retrieval hits for Hinglish phrasings of your most common questions |
| Speech synthesis | Replies in formal Hindi or pure English sound stiff, and English terms, brand names, acronyms and numbers inside Hindi speech can be mispronounced | Listen to replies containing product names, amounts and dates |
The language model layer needs its own testing. GLUECoS, a benchmark for code-switched text, found that, across both English-Hindi and English-Spanish, multilingual models given extra fine-tuning on code-switched text performed best on most of its tasks. Being multilingual helped, but targeted training helped more; that is evidence for testing and adapting, not a guarantee that fine-tuning solves code-switching. More recently, a January 2026 evaluation of LLMs on code-switched text, which introduced the CodeMixQA benchmark, reported “persistent challenges in both reasoning and generation under code-switching conditions” for the models it tested. So a clean transcript is necessary but not sufficient: test the reasoning layer on code-mixed input too.
Implementation Roadmap
- 1Collect real, consented calls from Hinglish-speaking callers and have native speakers transcribe them in your chosen script convention
- 2Score candidate speech-to-text models on that sample after normalising script, and shortlist those that handle code-mixing without a separate routing step
- 3Write the script policy into prompts, extraction fields and knowledge-base indexing so every system sees the same form of each word
- 4Test the language model on code-mixed transcripts and instruct the agent to mirror the caller's mix of Hindi and English
- 5Track p95 turn latency and repair turns per call after launch, and A/B test one change at a time
Designing Hinglish AI Agents That Hold Up
For callers in many parts of India, Hinglish is often the default rather than an edge case, so design for it from the start. Begin with your own audio: take a sample of real, consented calls, have native speakers transcribe them in the script convention you chose, and score candidate models on that set. Public benchmarks help with shortlisting; for example, the IndicVoices benchmark contains 7,348 hours of audio from 16,237 speakers across 22 languages. But a public benchmark doesn't know your product names, your plan terms or the way your customers shorten them.
Next, let the agent follow the caller. If the caller mixes languages, a reply in textbook Hindi sounds as out of place as a reply in formal English. Tell the model in the prompt to mirror the caller's mix and keep English product terms in English. Then listen to the voice: code-mixed text-to-speech has to pronounce product and brand names, acronyms, numbers and English terms inside Hindi sentences, and a voice that stumbles over “EMI” or an amount sounds unnatural even when every word is right. Finally, measure whole turns: track end-to-end turn latency at the 95th percentile (p95) together with repair turns per call, and treat a rise in either as a regression.
Zoice's own product builds in several of these choices. Its language support treats Hinglish as a dedicated part of the audio pipeline, not just language detection, with code-switching recognition that transcribes each span in the language it was spoken, and agents reply in the language the customer switches to. Knowledge-base chunks are tagged by language, and on AI phone calls each contact in a CSV campaign can be given their own language, across ten Indian languages plus English (India).
For measurement, Zoice's AI Summary & Analytics lets you run the test this post recommends. An experiment could pit a prompt that mirrors the caller's code-mixing against one that answers in formal Hindi, compare the two variants on conversion, cost and sentiment, and promote the winner. Each call also carries a voice-quality score weighted across latency, interruptions, silence and talk balance, p95 turn latency and per-message token and latency logs, and Ask Anything answers plain-English questions scoped by campaign, date range and call type.
Frequently Asked Questions
Does handling code-switching increase AI voice agent latency?
It depends on the design. A single model that transcribes mixed speech directly adds no separate routing step. A pipeline that first identifies the language and then routes audio to a monolingual recogniser adds a decision point, which can stall or reverse at every switch. Often the larger cost is a repair turn caused by misrecognition, because one extra exchange can easily outweigh savings inside the pipeline, so measure full turns rather than first-token time.
Can I just set the language to Hindi for Hinglish callers?
Only partly, since Hinglish isn't a separate language to select: English words get forced into Hindi spellings or misheard as Hindi words, which hurts names, numbers and product terms most. A model built for code-switching is a better start. For example, Deepgram lists Hindi among the ten languages supported by Nova-3 Multilingual, a model it describes as built for speech that mixes languages, and Sarvam states that Saaras V3 “handles code-mixed speech without distorting it”, describing its newer Saaras V4 as robust to code-mixing too. These are vendor capability claims, other providers offer code-switching models as well, and any of them should be checked on your own calls.
How accurate is speech-to-text for Hinglish today?
Published numbers vary widely and rarely compare cleanly. Sarvam reports about 19% WER for Saaras V3 on the IndicVoices benchmark (19.31% on the subset of its ten most popular languages), down from about 22% for Saaras V2.5. That is a V3 result; Saaras V4 is newer. In IndicVoices, 74% of the audio is extempore speech and 17% is conversational. It is a broad Indian-language benchmark, so the figure does not isolate code-mixed speech, and it is self-reported. Treat any single number as a reason to test, not a forecast.
Should Hinglish transcripts be written in Devanagari or Roman script?
Either can work. What matters is choosing one convention and enforcing it everywhere, evaluation references included. Whichever you pick, keep identifiers such as order IDs, email addresses and plan names in Roman script, and when you compare vendors, score every transcript in the same script.
If your callers switch languages mid-sentence, send us a few sample calls and we will score how an agent handles them, the same test this post recommends running on your own traffic.
Sources
- Universals and cultural variation in turn-taking in conversation — Proceedings of the National Academy of Sciences (via PubMed Central)
- Multilingual and code-switching ASR challenges for low resource Indian languages (MUCS 2021) — arXiv (Interspeech 2021)
- SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset — NeurIPS 2025 (arXiv)
- Hinglish voice AI speech recognition — Deepgram
- Building for Indian Languages — Sarvam AI Docs
- Introducing Saaras V3 — Sarvam AI
- Introducing Saaras V4 — Sarvam AI
- Nova-3 Multilingual: major WER improvements across languages — Deepgram
- GLUECoS: An Evaluation Benchmark for Code-Switched NLP — arXiv (ACL 2020)
- Can Large Language Models Understand, Reason About, and Generate Code-Switched Text? — arXiv (January 2026, CodeMixQA)
- IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages — arXiv (AI4Bharat)
Written by
Vicky Yadav
Backend & AI Systems Engineer, Zoice
Vicky Yadav is a Backend & AI Systems Engineer at Zoice, focused on building and scaling the engineering systems behind conversational AI products. He works across backend services, APIs, telephony integrations, real-time voice workflows, and production infrastructure, with a focus on making AI-powered systems reliable, scalable, and production-ready for real-world use cases.
Keep reading
All articles
AI Voice Agent Latency: What Happens in the Second After You Say “Hello”

Connect Plivo to Zoice: A Step-by-Step Guide to Putting an AI Agent on Your Phone Number

WhatsApp Business API Without a BSP: What Skipping the Middleman Actually Means
Ready to put an AI agent to work?
Deploy voice, WhatsApp, and chat agents across Indian languages — grounded in your knowledge and measured on every call.