Langoedge Blog
Voice AI Agent Latency: Why ElevenLabs' 100ms Model Still Means a 1.4-Second Call
Voice AI agent latency is the total delay a caller experiences between finishing a sentence and hearing the agent's reply — and it is almost never the same number a text-to-speech vendor puts on a model card. On September 28, 2026, ElevenLabs shipped Eleven v4 and v4 Turbo, its fastest and most expressive speech models yet, with a headline figure of roughly 150 milliseconds median time-to-first-sound. Independent, real-phone-call measurements of ElevenLabs' own voice agent stack put its actual latency closer to 1,424 milliseconds — nearly ten times higher. Neither number is wrong. They're measuring different things, and the gap between them is the single most misunderstood part of building a voice AI agent.
That gap is also, not coincidentally, exactly where a platform's telephony architecture either earns its keep or quietly taxes every call you run through it.
The voice AI agent latency number ElevenLabs just published
Eleven v4 is a real upgrade. According to TechCrunch's coverage of the launch, the new models support 90 languages (up from 70 in v3), add stacked expression tags for handling things like interruptions and escalations more naturally, and cut voice cloning down to 10 seconds of reference audio. Within hours, Artificial Analysis' Provider Voice Arena leaderboard put Eleven v4 in first place, with an Elo score around 1,315-1,319 from blind listener comparisons against Cartesia's Sonic 3.6 and Google's Gemini TTS models.
The latency claim is the part that matters for anyone building phone agents, not audiobooks. ElevenLabs' own product page states Eleven v4 Turbo has a "150ms median time to first sound," and trade coverage from Unite.AI breaks that down further: roughly 100ms of that is model inference, with the rest coming from network round-trip in a WebSocket streaming test. That's a genuinely fast model. It is not, and was never claimed by ElevenLabs to be, the latency a caller experiences on an actual phone call.
Voice AI agent latency measured on a lab WebSocket connection and voice AI agent latency measured on a real phone call are roughly a 10x gap for the same underlying model.
What voice AI agent latency actually measures on a phone call
A model benchmark times one thing: how fast the model can turn text into audio once it has the text. A phone call has to get through several more stages before that model ever runs, and every one of them adds milliseconds a leaderboard doesn't count.
Here's the actual sequence, on a live inbound call:
Six stages sit between "caller stops talking" and "caller hears a reply" — the TTS model that ElevenLabs' 150ms figure describes is only one of them.
The TTS step is stage five of six. Before it can even start, the platform needs to decide the caller has actually finished their turn (endpoint detection has to tolerate pauses, breaths, and "um"s without cutting people off), transcribe what they said, and run that transcript through an LLM that's deciding what to say back. In a real orchestration layer, that step is also updating state: did the caller ask for the wrong department, does a slot need re-checking, has this turn triggered a branch in the conversation graph. Daily's Pipecat team, in a February 2026 benchmarking writeup, put the rule of thumb at under 1,500 milliseconds total voice-to-voice latency for a conversation to feel natural, with roughly 700ms of that budget going to the LLM's time-to-first-token alone in a transcription-to-LLM-to-voice pipeline. A 100ms-faster TTS model helps. It does not close a 700ms LLM gap or a slow STT pass by itself.
Then there's the part model benchmarks structurally cannot see at all: getting the audio onto an actual telephone line. A WebSocket test runs over the open internet to a browser or a test harness. A real inbound call runs over SIP trunking into the PSTN, frequently bridged through a carrier, a media relay, and a jitter buffer that's smoothing for a lossy, codec-constrained circuit — not a browser tab. That's infrastructure a voice model provider doesn't operate and a model leaderboard was never measuring.
The real-call voice AI agent latency benchmark
This is exactly why an independent site called Openbenchmarks exists, and why its numbers matter more for evaluation than any vendor's model card. Its methodology: a caller robot dials each platform's live phone agent, both sides of the call are recorded on a single clock, and speech endpoints are detected from the actual audio using Silero VAD rather than trusting any platform's self-reported timestamps. The metric is TTFAB — time to first audio byte, measured from when the caller stops speaking to when the agent's reply audio begins.
Here's what that produced across five platforms, as median TTFAB from real phone calls:
Every platform in this voice AI agent latency benchmark lands between 1.3 and 1.75 seconds median TTFAB on real calls — nowhere close to any vendor's WebSocket demo number, ElevenLabs included.
Bland AI's p95 stretches to 2,248ms; Retell AI's median alone is already 1,740ms. The full range across all five platforms — 1,296ms to 1,740ms median, per Openbenchmarks' published data — is a narrower band than most marketing pages would suggest, and every single one of them is well north of what a 100-150ms TTS benchmark implies is possible.
Claim: the fastest platform in a real-call voice AI agent latency benchmark should be the one with the best speech model.
Evidence: it isn't. Telnyx sits at the top of the Openbenchmarks list at 1,296ms median TTFAB, and Telnyx is a telephony carrier — it doesn't sell a proprietary TTS model at all. The lowest latency in this benchmark belongs to the infrastructure layer, not the speech model layer.
Why direct SIP trunking is a bigger lever than the TTS model
If TTS synthesis is one stage out of six, and every platform in the Openbenchmarks list already uses a fast enough model, the remaining latency has to be coming from somewhere structural: how many hops the audio takes between the caller's phone and wherever inference is actually running, and how much of the pipeline is chained sequentially versus run concurrently.
This is the specific problem we built Langoedge's telephony layer around. Rather than routing calls through an intermediary voice API and re-bridging into LiveKit, Langoedge connects directly via Telnyx and Twilio SIP trunking, cutting out a hop most stacked voice-AI-on-top-of-voice-API architectures carry by default. Our own internal design target for that path is sub-1.2 second end-to-end latency. We're stating that plainly as our target figure, not as an independently audited number in the way Openbenchmarks measures its five platforms, because we haven't submitted to that specific benchmark yet. It's the same reason we're flagging it here rather than just quoting it as fact: a number a vendor states about itself and a number a third party measures from real calls are different kinds of evidence, and conflating them is the exact mistake this whole ElevenLabs model-card-versus-phone-call story is about.
Claim: heavier backend work — a booking check, a CRM write, a database lookup — doesn't affect voice AI agent latency as long as the speech model is fast.
Evidence: it does, directly, if that work runs synchronously inside the same voice loop that's also waiting on STT and TTS. Whatever the backend call takes gets added straight to TTFAB, regardless of how fast the TTS model is. Keeping that logic off to the side, in a process the voice loop doesn't block on, is a structural choice that shows up in latency numbers whether or not a vendor frames it that way.
What to ask a voice AI platform about latency before you build on it
For a developer or agency evaluating platforms — which is most of the audience actually reading vendor comparison pages — the practical takeaway from the Eleven v4 launch isn't "ElevenLabs is fast" or "ElevenLabs is slow." It's that a TTS model benchmark and a voice AI agent latency number answer different questions, and a vendor is under no obligation to volunteer which one they're quoting.
Three questions cut through most of the marketing noise:
Is this a model benchmark or a phone-call benchmark? A number measured over WebSocket to a browser, or against a text prompt in isolation, tells you about the model. A number measured from a live inbound PSTN call tells you about the product you'd actually be deploying.
Who measured it? A vendor's own dashboard and an independent benchmark using the same recorded-audio methodology across competitors are not equivalent evidence, even when they use the same unit (milliseconds) and the same rough definition (time to first response).
Does the number include backend work, or just the happy path? A demo script with no database lookup and no branching logic will always be faster than the agent you actually ship, which usually needs both.
FAQ
What is voice AI agent latency?
Voice AI agent latency is the time between a caller finishing what they're saying and hearing the agent's spoken reply begin, measured end-to-end across speech-to-text, LLM reasoning, text-to-speech, and the telephony network — not just the time any single model takes to generate audio.
Why does ElevenLabs report 100-150ms latency if real calls take over a second?
ElevenLabs' published figure describes Eleven v4 Turbo's own processing time, measured largely in a WebSocket streaming test — it is a model-level number. Independent phone-call testing from Openbenchmarks measured ElevenLabs' full voice agent stack at 1,424ms median TTFAB, because a real call also includes endpoint detection, transcription, LLM reasoning, and delivery over SIP/PSTN, none of which the model-level figure covers.
What is TTFAB and why does it matter more than a model benchmark for evaluating voice AI platforms?
TTFAB stands for time to first audio byte: the measured gap between a caller going silent and the agent's reply audio starting, captured from actual recorded phone calls rather than vendor-reported timestamps. It matters more than a model benchmark because it's the number that reflects what a caller on a real phone line experiences, including every pipeline stage a model card doesn't measure.
Does a faster text-to-speech model make a voice AI agent noticeably faster overall?
A faster text-to-speech model helps, but it's one stage in a five-or-six-stage voice AI pipeline, and speech-to-text plus LLM reasoning typically account for a larger share of total latency — industry benchmarking has put LLM time-to-first-token alone at roughly 700ms within a sub-1,500ms voice-to-voice budget. A 50ms improvement in TTS inference is real but rarely the deciding factor in whether a voice AI agent feels responsive on a phone call.
How does Langoedge's architecture address voice AI agent latency?
Langoedge connects calls directly through Telnyx and Twilio SIP trunking rather than bridging through an intermediary voice API layer, and keeps heavier backend logic like database writes or CRM calls running asynchronously alongside the voice loop instead of blocking it. Our internal design target for that path is sub-1.2 second end-to-end latency; this is a stated target rather than a number independently benchmarked the way Openbenchmarks measures TTFAB.
Sources
- ElevenLabs' new v4 speech model supports more expression control and 90 languages — TechCrunch, September 28, 2026
- Eleven v4 and Eleven v4 Turbo — official ElevenLabs product page
- Text to Speech Leaderboard — Artificial Analysis
- ElevenLabs Launches Eleven v4 With Low-Latency Turbo Variant — Unite.AI
- Voice agent latency benchmark — TTFAB measured from real phone calls, Openbenchmarks
- Benchmarking LLMs for Voice Agent Use Cases — Daily