Back to blog
Voice AI6 min read

Voice AI Latency: What Mercury Voice's 320ms Means for Calls

Inception's Mercury Voice reports 320ms to first token. Here is where voice AI latency really comes from and how to cut it in your own build.

HM
Harshit Makraria
October 3, 2026

We've spent the last 11 months shipping voice agent deployments for coaches, consultants, fintech, real estate, and a handful of edge cases. Ninety-six in production. Here's what we've learned about what actually works in 2026.

1. The model isn't the bottleneck anymore

GPT-4o-realtime, Claude 3.5 Sonnet voice, and the open-source equivalents are good enough for 92% of production scenarios. Telephony latency, audio processing pipelines, and prompt routing are now the failure modes not LLM quality.

If your agent feels janky, audit your audio path before you audit your prompts. Eight times out of ten, that's where the friction lives.

"The agents that work feel like infrastructure. The agents that fail feel like party tricks."

2. Voice ≠ chatbot with audio

Every team that tries to port their chatbot prompt to voice fails the same way: too verbose, too formal, too explainer-y. Voice is improv. You need shorter turns, callback handles, and graceful interruption.

3. The handoff is the product

The best voice agent in the world is useless if the post-call sync is broken. Notes go to CRM. CRM triggers sequence. Sequence books follow-up. Calendar invites human. That is the system. The voice piece is one component.

If you want to see a live example, our AI calling system is running in production for loan servicing and collections you can see the real numbers on the case studies page.

Pick up a phone call and wait one full second before the other person answers. You already feel something is off. By the second second, you are saying "hello?" That is the exact problem every business voice agent fights, and voice AI latency is now the number that separates a system customers tolerate from one they hang up on. This week the race got faster. Inception introduced Mercury Voice, a diffusion language model built for voice agents, with a reported median time to first answer token of 320 milliseconds. Microsoft also unveiled new voice models, including a real-time streaming speech-to-text model called MAI-Transcribe-2-Streaming. If you run or plan to run phone automation, here is what these numbers mean and how to get your own latency down.

Why Voice AI Latency Decides Whether Calls Work

Humans take turns in conversation with a gap of roughly 200 to 300 milliseconds. Anything much past half a second starts to feel like a bad international line. Past a second, callers interrupt, repeat themselves, or assume the bot is broken. Interruptions then trigger more bugs, because the agent is still talking about an old answer while the caller has moved on.

This is why latency is not a vanity metric. It drives three business outcomes directly:

  • Completion rate: callers who feel lag drop off before the agent finishes qualifying or booking.
  • Trust: a fast, natural reply reads as competent. A slow one reads as a robot, even when the answer is correct.
  • Cost per call: long silences and repeated turns inflate minutes, and minutes are what you pay for.

Where the Milliseconds Actually Go

A voice agent is a pipeline, and each stage adds delay. Understanding the stack tells you where a faster model helps and where it does not.

  • Telephony and network: carrier hops and audio transport add delay before your software hears a word. Region choice matters here.
  • Speech to text: the system must decide the caller finished speaking, then transcribe. Streaming transcription starts working while the caller is still talking, which is why models like Microsoft's new real-time option matter.
  • The language model: time to first token is the critical figure. The reasoning model can take seconds if it is large or is calling tools. This is the stage Mercury Voice targets.
  • Tools and data lookups: a CRM query or calendar check can add several hundred milliseconds on its own.
  • Text to speech: the voice must start speaking before the full sentence is generated, or the caller waits for the entire reply.

Add those up and a naive build lands well above one second. A tuned build keeps the whole loop under about 800 milliseconds, and the best ones feel instant. A 320 millisecond first token is a real step, but it is one stage out of five. You still have to engineer the rest.

Why Diffusion Language Models Matter for Voice

Most language models generate one token at a time, left to right. Diffusion language models refine many tokens in parallel, which can make output start and finish faster. For long documents that is a speed story. For voice, the interesting part is the very short, conversational reply, where time to first token dominates the experience. The honest caveat: a faster first token does not help if the answer is wrong, and independent testing on your own call types matters more than a vendor figure. Treat the 320 millisecond number as reported, not as a guarantee for your workload.

The strategic point is more durable than any single model. Speech and language components are getting cheaper and faster every quarter, so your architecture should let you swap them without a rebuild. If your voice agent is hard-wired to one provider, every launch like this one is a migration project instead of a config change.

How to Cut Voice AI Latency in Your Own Build

You do not need to wait for new models to fix a slow agent. These are the changes that usually move the number most:

  • Stream everything. Stream transcription in, tokens out, and audio out. Never wait for a full sentence at any stage.
  • Use a small, fast model for the live turn. Route simple conversational replies to a quick model and reserve the large model for hard reasoning that can happen off the critical path.
  • Prefetch data. Load the caller's record the moment the call connects, using caller ID, so no lookup happens mid-sentence.
  • Mask unavoidable delays. If a tool call takes time, have the agent say a short natural phrase while it works, the way a person would say "let me check that."
  • Co-locate services. Keep telephony, model, and voice in the same region to avoid cross-ocean round trips.
  • Measure per stage. Log time for each hop on every call. You cannot fix a total you have not broken down.

At Nexica AI we build voice systems with this stage-by-stage discipline. Our work has handled $48.9M in accounts, and calls like that only stay on the line when the agent responds like a person. We also build TCPA compliant flows, because speed means nothing if the call should never have been placed.

What to Do This Week

First, time your current agent. Place ten real test calls and record the gap between the caller finishing and the agent speaking. If the median is above one second, you have a revenue leak, not a polish problem. Second, map which stage owns the biggest share and fix that one first. Third, make your stack model-agnostic so faster speech and language models, like the ones announced this week, can be dropped in without touching your logic. Finally, set a latency budget per stage and alert when a call exceeds it, the same way you would for uptime.

The companies that win phone automation in 2026 will not be the ones with the flashiest voice. They will be the ones that answer before the caller has time to doubt. See how this fits into your wider workflow automation and our case studies for real numbers.

If you want this built for your business, book a 20-minute call with Nexica AI. We build production-grade AI systems in 14 days.

AI CallingVAPIProductionPlaybook
Want this built for your business?See our AI calling system
Free AI Audit