AI voice agent latency

AI Voice Agent Latency: What Happens in the Second After You Say “Hello”

SB

Shivansh Bhardwaj

AI Product & Systems Engineer

October 6, 202613 min read
AI Voice Agent Latency: What Happens in the Second After You Say “Hello”

You say “Hello” into the phone. For a moment nothing happens, and then the agent answers. If that moment is short, the call feels like a conversation. If it drags on, the caller says “Hello?” again, the agent starts answering the first greeting just as the second one arrives, and both sides talk over each other. When people complain about AI voice agent latency, they usually mean this one gap: the time between the caller finishing a sentence and hearing the first syllable of the reply.

That gap is several delays added together. A turn passes through audio capture, network transport, voice activity detection, transcription, end-of-turn detection, reply generation, speech synthesis and playback. Some of the time is compute, some is network, and some is deliberate waiting. You can't shrink a number until you know what it's made of, so this post follows one spoken turn through the voice AI pipeline, shows where the time typically goes, and then covers the techniques that actually shorten it.

The short version

  • Stream every stage, so the first words play before the full reply exists.
  • End turns on meaning, not only on a silence timer.
  • Host speech-to-text, the LLM and text-to-speech close together and near your callers.
  • Measure p95 voice-to-voice latency per turn, not the average.

Why AI Voice Agent Latency Breaks Conversations

Human conversation runs on tight timing. Tanya Stivers and colleagues filmed informal conversations in ten languages and timed how quickly people answered yes/no questions, from the end of the question to the start of the answer. Across their whole dataset the average gap was about 208 ms and the median about 100 ms. Each language had its own tempo, but every language's average stayed within roughly 250 ms of that overall figure (Stivers et al., PNAS, 2009). It is reasonable to expect callers to bring similar expectations to a call answered by software.

Voice agents start behind. A cascaded agent usually waits for the turn to close and then runs several models one after another. According to the Pipecat team's illustrated primer on voice agents, a fairly typical cloud round trip from the caller's microphone to the agent and back comes to about 1,200 ms, and the primer names 1,500 ms as the voice-to-voice figure conversational applications should aim for. Treat that as a target to reach in production, not as what feels natural between two people.

Two things follow. First, the slow turns matter more than the typical one, so a good median can hide a poor experience; the endpointing section below shows why. Second, latency and turn-taking errors trade off against each other. An agent tuned to reply as fast as possible will cut people off mid-thought, and one tuned never to interrupt will leave awkward silences. Most latency work is about handling that trade-off better, not just about faster models.

Key Insight

Most of an AI voice agent's delay comes from deciding that the caller has finished and waiting for the language model's first token, which together take 950 of 1,293 ms in Pipecat's typical cloud breakdown.

On this page

How AI Voice Agents Work, Stage by Stage

A cascaded voice agent has separate components for hearing, thinking and speaking, connected by streams. Here is what happens to the word “Hello” on its way through, and where each stage adds time.

Hearing: capture, transport and transcription

The caller's microphone or handset digitises the audio, an encoder compresses it, and it travels to the agent's servers. Browser and app calls usually go over WebRTC; phone calls cross the telephone network and a SIP trunk. Every hop adds encoding time, network transit and a jitter buffer, which holds back a little audio so that late or out-of-order packets still play smoothly. The reply makes the same journey in reverse, so each of these costs is paid twice. In Pipecat's breakdown, the audio path in both directions comes to roughly 200 ms.

On the server, a voice activity detector (VAD) decides, frame by frame, whether the caller is talking. This step is cheap. The project page for Silero VAD, an open-source detector, says it needs under a millisecond of single-thread CPU time to classify a chunk of 30 ms or longer. Speech frames then stream into a speech-to-text (STT) model, which sends partial transcripts while the caller is still talking and a final transcript once the utterance is complete.

Deciding: has the caller finished?

The agent can't reply until it believes the turn is over, and silence is an ambiguous signal: callers pause to recall an account number, to think, or to breathe. The simplest rule ends the turn after a fixed stretch of silence. In LiveKit's agent framework, for example, that wait is set by a minimum endpointing delay that is 500 ms unless you change it. That time is spent on purpose, before any model has started on the answer.

Better endpointing reads the words as well as the silence. LiveKit has described an end-of-utterance model built as a small transformer of about 135 million parameters, based on Hugging Face's SmolLM v2 so it could run on CPU, with inference of roughly 50 ms (vendor figure). Its multilingual model v0.4.1-intl, built on a different base model, produced 39.23% fewer false-positive interruptions than v0.3.0-intl in LiveKit's benchmarks, without adding response latency (vendor figure). “My number is nine eight seven…” followed by a pause is plainly unfinished, while “Yes, that's right.” is plainly done. A silence timer can't tell them apart; a language model can.

Answering: the LLM, TTS and playback

Once the turn is closed, the transcript and conversation history go to the LLM. For latency, the number that counts is time to first token, not total generation time, because a streaming pipeline can start speaking long before the reply is complete. The first sentence goes to text-to-speech, and its time to first audio depends heavily on the voice model. ElevenLabs' model list gives about 75 ms for Flash v2.5 and about 280 ms for the more expressive Eleven v3 Conversational (vendor figures). The audio then returns to the caller through another encoder, network hop and jitter buffer.

Each of these stages is a separate component that you can measure, swap and tune. The core components below summarise what each one does and where its latency comes from.

Core Components

1Capture and Transport

The caller's audio is digitised, encoded, sent over WebRTC or a telephone and SIP path, and smoothed by jitter buffers. The reply makes the same trip back, so every hop is paid twice.

2Voice Activity Detection

A lightweight classifier labels short audio frames as speech or non-speech. It decides what reaches the transcriber and gives the first signal that the caller may have stopped talking.

3Speech-to-Text and End-of-Turn Detection

Streaming STT emits partial and final transcripts while an endpointing step decides that the turn is over, using either a fixed silence threshold or a semantic turn model. Much of the waiting in a turn happens at this step.

4LLM Response Generation

The transcript and conversation history go to the language model. Time to first token matters more than total generation time, and it depends on model size, prompt length and caching, tool calls and how far away the model is hosted.

5Text-to-Speech and Playback

The first sentence of the reply is synthesised and streamed back as audio while the rest is still being generated. Choosing a fast voice model or a more expressive one changes the time to first audio byte.

Where the Milliseconds Go in a Voice AI Pipeline

Pipecat's primer publishes a stage-by-stage breakdown of one typical cloud round trip; the table below regroups its figures into four buckets. Treat it as one illustrative budget, not a benchmark.

BucketTime in Pipecat's breakdownWhat moves it
Deciding the turn is over (transcription and endpointing)300 msSTT speed, silence threshold or turn model
First token from the LLM650 msModel size, prompt length and caching, provider load, hosting distance
First audio from TTS120 msVoice model, streaming support
Audio path both ways, plus 20 ms to assemble the first sentence for TTS223 msTransport, codec, jitter buffers, distance to the caller
Voice-to-voice total1,293 msEverything above, end to end

In Pipecat's figures, the three model buckets add up to 1,070 ms, and the remaining 223 ms (1,293 − 1,070) is the audio path plus sentence assembly. The “about 1,200 ms” quoted earlier is Pipecat's own rounded description of this same trip. You will also see different endpointing figures in this post, and they don't contradict each other: Pipecat's 300 ms is a budget for one well-tuned turn, LiveKit's 500 ms is a library default for the silence wait, and Deepgram's 1.5 s (below) is a 95th-percentile figure for the slowest turns.

One turn, step by step (illustrative)

Laying Pipecat's figures out as a timeline shows what the caller experiences. These are first-token and first-byte times, so the example already assumes streaming. For simplicity it counts the model stages first and adds the audio path, in both directions, as one step at the end.

  1. Time zero: the caller finishes saying “Hello”.
  2. ~300 ms: the transcript is final and the turn is declared over.
  3. ~950 ms: the LLM's first token arrives.
  4. ~1,090 ms: the first sentence has been assembled and TTS returns its first audio.
  5. ~1,293 ms: after about 203 ms of audio path in both directions (capture, encoding, network, jitter buffers, decoding and playback), the caller hears the reply.

In this example, the caller hears nothing for roughly 1.3 seconds. Two buckets dominate: the LLM's first token and the decision that the turn is over, which together take 950 ms of the budget. A pipeline that waits for complete transcripts, complete replies or complete audio instead of streaming adds each stage's full processing time on top.

The endpointing tail

Averages hide the worst turns. Deepgram says Flux, its speech recognition model with built-in turn detection, typically reaches 1 s at p90 and 1.5 s at p95 for end-of-turn latency (vendor figures). By Deepgram's own figures, about one turn in twenty takes 1.5 s or more just to detect end of turn. A dashboard that shows only the average never surfaces those turns, and they are the ones callers remember, which is why p95 is the number to track.

Cascaded pipelines vs speech-to-speech

The alternative to a chain of models is one model that listens and speaks directly. Kyutai's Moshi paper (2024) argues that the spoken-dialogue pipelines its authors set out to replace, with VAD, speech recognition, a text model and TTS chained together, added several seconds of latency between turns. For its own full-duplex speech model it reports 160 ms in theory and 200 ms in practice. Two caveats keep this fair. The “several seconds” describes the pipelines the authors were comparing against, not a tuned modern cascade like the 1.3-second budget above. And Moshi's figures are model-level, before any network or telephony leg a phone call would add. The numbers aren't like for like, but the underlying point holds: every hand-off between models costs time.

Speech-to-speech is available commercially too. According to OpenAI's API changelog, its Realtime API for speech-to-speech launched on 1 October 2024 and became generally available on 28 August 2025. The same changelog lists GPT-Live 1, for full-duplex voice conversations that continue while a backend model handles reasoning, as generally available from 10 September 2026.

The trade-off is control. A cascade gives you a text layer at every stage that is easy to inspect, ground, redact and audit, and you choose each model independently; with speech-to-speech, those controls depend on what the provider exposes. The cascade pays for that flexibility with extra hand-offs, which is where the techniques below focus.

Implementation Roadmap

  1. 1Instrument every turn: log end-of-turn time, LLM time to first token, TTS time to first byte and total voice-to-voice latency, and track p95 alongside the median
  2. 2Stream every stage (partial transcripts, LLM token streaming and sentence-by-sentence TTS) so the first words play before the full reply is written
  3. 3Replace a fixed silence threshold with semantic end-of-turn detection, and tune it against both response latency and the rate of false interruptions
  4. 4Cut LLM time to first token: keep a stable prompt prefix so provider caching applies, trim history, size the model for conversational turns, and speak a short acknowledgement before slow tool calls
  5. 5Host STT, LLM and TTS close to each other and to your callers, choose a fast streaming voice, then re-measure p95 after each single change

How to Reduce AI Voice Agent Latency

None of these techniques is exotic; the gains come from applying them consistently.

Start by streaming every stage so the stages overlap. STT should send partial transcripts, the LLM should stream tokens, and TTS should start on the first sentence or clause instead of waiting for the full reply. The caller then waits only for that first clause to be generated and synthesised, so what matters is the length of the first clause, not the length of the whole answer. Asking the model to open with a short sentence can help.

Next, make endpointing smarter rather than just shorter. Lowering the silence threshold cuts latency but makes interruptions more likely; a semantic turn model can wait through unfinished sentences and answer finished ones quickly. Deepgram claims that Flux can shorten the time an agent takes to reply by 200–600 ms relative to pipeline approaches (vendor figure).

Speculative endpointing goes a step further: you start the LLM on a likely end of turn and discard the result if the caller keeps talking. Flux exposes this through its eager_eot_threshold setting. Deepgram's Flux page says that at a threshold of 0.3–0.5 the early EagerEndOfTurn event typically fires 150–250 ms ahead of the final EndOfTurn, in exchange for making 50–70% more LLM calls (vendor figures).

Hosting matters next. The Pipecat team reports agents reaching 500 ms voice-to-voice by running every model in one GPU cluster and tuning for latency over throughput. Even without your own GPUs, keeping STT, the LLM and TTS in one region near your callers removes round trips from every turn.

On the LLM side, prompt size is the lever you control most directly, and prompt caching is the standard fix. Put the parts that never change, such as system instructions, tool definitions and fixed examples, at the very start of the prompt and keep them identical from turn to turn. Then append what changes: conversation history, tool results and retrieved passages. Providers such as Anthropic and OpenAI can then reuse a repeated prompt prefix instead of reprocessing it, which their documentation says reduces latency as well as cost, though only once the reusable prefix passes a model-specific minimum length (512 to 4,096 tokens on Anthropic's models, 1,024 tokens on OpenAI's newer ones). Beyond caching, trim history, use a model sized for conversational turns, and cover slow tool calls with a short spoken acknowledgement such as “Let me check that”.

What to check first

When an agent feels slow, split each slow turn into end-of-turn time, LLM time to first token, TTS time to first audio and transport, then work on the largest share:

  1. If end-of-turn time dominates, review the silence threshold and try semantic turn detection, watching the interruption rate as you change it.
  2. If the LLM's first token dominates, check prompt size and caching, model size and region, and whether tool calls run before the first words.
  3. If TTS first audio dominates, compare voice models, since more expressive ones can be slower.
  4. If slow turns cluster on certain routes or carriers, check hosting distance and codec transcoding.

That breakdown needs per-turn data. On Zoice, each call gets a voice-quality score weighted across latency, interruptions, silence and talk balance, alongside p95 turn latency, and tokens and latency are logged per message, so call analytics show which turn slowed down. If you bring your own carrier, SIP trunks let you set a codec preference order to match your carrier, and each trunk has a test button that sends a real SIP OPTIONS probe. The roadmap above puts these steps in the order a team would usually tackle them.

A note on the numbers: every figure in this post is typical, vendor-reported or illustrative. Your results will depend on your models, hosting, carrier path and configuration, so measure before and after each change.

Frequently Asked Questions

What is a good AI voice agent latency?

There is no universal number. Pipecat's primer names 1,500 ms voice-to-voice as a target for conversational applications, and the Pipecat team reports reaching 500 ms with co-located models tuned for latency. Set your own target on p95 rather than the average, because callers notice the slow turns.

Why does my agent interrupt callers when I make it faster?

Usually because the silence threshold was shortened. A shorter wait closes the turn sooner, but it also mistakes mid-sentence pauses for a finished reply. Semantic end-of-turn detection handles this better than a timer because it reads the words, so tune latency and interruption rate together.

Does the phone network add latency compared with app calls?

Yes, though how much depends on the route. A phone call adds carrier hops, SIP signalling, codec encoding and possibly transcoding, and jitter buffering in both directions. You can't remove the carrier path, but you can host the agent close to where calls enter your infrastructure and avoid needless codec conversions.

Should I switch to a speech-to-speech model to fix latency?

Only if a tuned cascade still falls short. A cascade keeps a text layer you can use for knowledge-base grounding, redaction, structured extraction and auditing, and lets you choose each model. Streaming, better endpointing, prompt caching and hosting fix many latency problems without changing the architecture, so measure first.

To see where the time goes on your own call flows, explore AI phone calls on Zoice or talk to our team.

Sources

  1. Voice AI & Voice Agents: An Illustrated Primer — Daily (Pipecat) — Kwindla Hultman Kramer
  2. Universals and cultural variation in turn-taking in conversation (Stivers et al., 2009) — Proceedings of the National Academy of Sciences
  3. snakers4/silero-vad — Silero (GitHub)
  4. Using a transformer to improve end of turn detection — LiveKit
  5. Improved end-of-turn model cuts voice AI interruptions 39% — LiveKit
  6. Models — ElevenLabs
  7. Flux – Conversational Speech Recognition — Deepgram
  8. End-of-Turn Detection Parameters (Flux) — Deepgram documentation
  9. Moshi: a speech-text foundation model for real-time dialogue — Kyutai (arXiv)
  10. Changelog — OpenAI
  11. Prompt caching — Anthropic documentation
  12. Prompt caching — OpenAI documentation
SB

Written by

Shivansh Bhardwaj

AI Product & Systems Engineer

Shivansh Bhardwaj is an AI Product & Systems Engineer at Zoice, focused on building and scaling production-grade AI products and conversational systems. His work spans AI agent architecture, real-time voice pipelines, LLM orchestration, latency optimisation, and multilingual AI experiences. He writes about emerging AI technologies, engineering challenges, and practical approaches to building fast, reliable, and scalable AI products for real-world business use cases.

AI ProductsConversational AIAI AgentsLLM EngineeringVoice AI

Ready to put an AI agent to work?

Deploy voice, WhatsApp, and chat agents across Indian languages — grounded in your knowledge and measured on every call.