The defining constraint of conversational voice artificial intelligence is latency. Human conversational turn-taking involves remarkably low latency, with English speakers averaging a mere 239-millisecond delay between utterances, and Japanese speakers averaging an astonishing 7-millisecond delay1. When an AI agent exceeds this deeply ingrained biological and cultural expectation, the interaction ceases to feel like a conversation and instead feels like a rigid, walkie-talkie-style transaction. The psychological impact of latency is profound: when the delay stretches beyond the roughly 700-millisecond threshold, users experience cognitive dissonance, assuming the agent has failed to hear them, leading to overlapping speech, repeated commands, and ultimate abandonment of the interface.
Historically, standard serial dialogue architectures have struggled immensely with this constraint, often imposing overall response latencies of 1000 to 1500 milliseconds1. To compensate for these severe delays, engineers have deployed various stopgap measures, such as using large language models (LLMs) like GPT-4 to predictively fill in missing context from dropped words at the end of an incomplete utterance with over 60% effectiveness1. However, predictive compensation cannot fully mask the underlying physical delay of the system. The challenge is exacerbated when these systems are deployed over traditional telecommunications networks. Modern voice AI platforms add between 400 and 620 milliseconds of end-to-end processing delay per conversational turn, and that figure rides on top of network transit that is itself budgeted at just 150 milliseconds one-way before human-to-human Voice over IP (VoIP) call quality begins to tangibly degrade2.
Achieving sub-second latency is therefore not merely an academic benchmark or a vanity metric; it is the fundamental engineering prerequisite for deploying viable, commercial voice interfaces. The difference between a 600-millisecond response and a 1200-millisecond response is the difference between an agent that is perceived as “human” and one that is perceived as “broken.” This architectural challenge requires meticulous optimization across the entire stack, forcing engineering teams to carefully evaluate the trade-offs between traditional chained pipelines and emerging native speech-to-speech models, while simultaneously navigating the hostile latency environments of public telephony networks.
The engineering landscape for real-time voice agents is currently divided between two dominant architectural paradigms: the traditional chained pipeline and the emerging native speech-to-speech model. The chained pipeline operates by stitching together three distinct, specialized components. First, a Speech-to-Text (STT) model transcribes the incoming audio into text; second, an LLM processes that text to generate a textual response; and third, a Text-to-Speech (TTS) model synthesizes the text back into audio. While highly modular, this architecture forces serialization. The LLM cannot begin generating meaningful contextual responses until the STT has finalized a chunk of its output, and the TTS cannot synthesize natural prosody until a sufficient chunk of text is generated by the LLM. This inherent serialization is the primary culprit behind the 1000 to 1500 millisecond latencies historically observed in voice AI systems.
In stark contrast, native speech-to-speech models bypass the intermediate text bottleneck entirely. OpenAI’s GPT-4o exemplifies this shift, replacing the previous voice mode’s three-model pipeline with a unified end-to-end neural network. This architectural consolidation reduces average response latency to 320 milliseconds and lowers API costs by 50% compared to GPT-43. By utilizing a natively multimodal architecture, GPT-4o can process and respond to audio inputs in as little as 232 milliseconds4. Google’s Gemini 1.5 Pro takes a related but distinct approach: a natively multimodal architecture that ingests audio directly, processing up to 107 hours of audio (9.7 million tokens) without segmentation and achieving a 5.5% Word Error Rate (WER) on 15-minute videos5. Gemini natively understands audio content alongside text and images, providing a more holistic reading of the user’s intent — though unlike GPT-4o it responds in text rather than speech, so a voice agent built on it still requires a separate synthesis stage6.
Table 1: Comparison of natively multimodal audio model benchmarks
| Model | Architecture Type | Reported Latency / Speed | FLEURS Benchmark Score (WER) |
|---|---|---|---|
| OpenAI GPT-4o | Unified End-to-End Neural Network | 232ms – 320ms response time | 5.4 |
| Google Gemini 1.5 Pro | Natively Multimodal (audio input, text output) | Processes up to 107 hours unsegmented | 6.6 |
Evaluating these specific components provides the empirical baseline necessary to contrast against the human-perception threshold. In published comparisons on the FLEURS automatic speech recognition benchmark, which measures word error rates where lower is better, GPT-4o is reported at 5.4 against Gemini 1.5 Pro’s 6.67. However, the choice between chained and native architectures is not strictly binary based on speed alone. Chained pipelines offer unparalleled modularity, allowing engineering teams to swap in specialized STT models for specific accents or highly technical domain vocabularies, and to cache text responses for compliance auditing. Native models, while offering superior raw speed and emotional nuance by processing acoustic features directly, often lock developers into a single ecosystem and present challenges for strict textual logging and deterministic reasoning pathways. For agents whose actual job is transactional — writing bookings, orders, and records into external systems — the distinction is decisive: text-native models remain measurably more reliable at producing correct structured tool calls than audio-native ones, and in a transactional domain a small error rate on the write path is not a rounding error but a missed appointment or a corrupted record. The chained pipeline also produces a text transcript as a byproduct of normal operation, and that transcript is precisely the artifact that audit, quality assurance, and stage-by-stage failure attribution require.
To engineer a system that consistently operates under the 700-millisecond threshold, teams must construct a precise latency budget, isolating and measuring the delay introduced at every stage of the conversational turn. The total latency of a voice AI round trip is the sum of multiple sequential and overlapping processes. The first critical stage is Voice Activity Detection (VAD) and turn-taking, which determines when the user has finished speaking. Following this is the STT processing time, often measured as Time to First Token (TTFT). Next is the LLM inference time, and finally, the TTS synthesis, measured as Time to First Byte (TTFB). All of this is enveloped by the inescapable realities of network round-trip times.
Ltotal = tVAD + tSTT + tLLM + tTTS + tnet
When examining the empirical data from production chained pipelines, the margins for error are razor-thin. In a 2026 Coval benchmark of 2,400 runs, Deepgram Nova 3 achieved a median TTFT of 992 milliseconds, while ElevenLabs Scribe v2 recorded a slower median TTFT of 2,080 milliseconds8. However, ElevenLabs Scribe v2 Realtime utilizes predictive streaming for “negative latency” to achieve sub-150ms latency, whereas Deepgram’s Nova-3 provides sub-300ms streaming latency, with integrated end-of-turn (EOT) detection at approximately 260 milliseconds supplied by its companion Flux turn-taking model9. Consistency is just as vital as raw speed; despite having a slower absolute TTFT in the Coval benchmark, ElevenLabs Scribe v2 demonstrated the most consistent latency distribution with a standard deviation of 60 milliseconds, compared to Deepgram Nova 3’s 125 millisecond standard deviation8.
On the synthesis side, the TTS first-audio metrics often reveal a significant gap between laboratory benchmarks and production reality. ElevenLabs’ Flash v2.5 model claims inference speeds of approximately 75 milliseconds, but production benchmarks published by Deepgram — a competing vendor, citing Podcastle’s measurements — put the median TTFB at 255 milliseconds, a 180-millisecond gap attributable to network and application overhead10. The same source reports Deepgram’s own Aura-2 text-to-speech model achieving a sub-200ms median TTFB10. If an architecture relies on a 260ms VAD trigger, a 300ms STT/LLM generation phase, and a 255ms TTS synthesis phase, the system is already operating at 815 milliseconds, failing the human perception test before network transit is even fully accounted for. This strict anatomical breakdown demonstrates why concurrent execution and aggressive streaming are not optional enhancements, but absolute requirements for chained pipelines.

Figure 1: The per-turn latency budget of a chained pipeline. The worked example (260 + 300 + 255 ms) exceeds the ~700 ms human-perception threshold before network transit is counted.
While native speech-to-speech models and optimized chained pipelines excel in controlled web-audio environments using modern protocols, these advantages often evaporate when agents are deployed over the Public Switched Telephone Network (PSTN). The transition from web to telephony introduces severe bandwidth constraints and legacy hardware interventions that fundamentally degrade both speed and accuracy. Using SIP for PSTN integration typically restricts voice calls to the narrowband 8kHz G.711 codec, which meaningfully degrades STT accuracy compared to the wideband Opus codec natively supported by WebRTC11. This loss of acoustic fidelity forces the LLM to work with flawed transcriptions, increasing hallucination rates and requiring additional conversational turns to clarify user intent, thereby ballooning the effective latency of the interaction.
Furthermore, the physical infrastructure of telecom carriers introduces unavoidable processing delays. Hardware DSP-based transcoding between modern codecs like Opus and legacy codecs like G.711 on Session Border Controllers (SBCs) adds approximately 20 milliseconds of latency per call leg12. Because carrier-grade SBC transcoding is frequently bound to dedicated hardware DSPs, many operators choose to pass G.711 end-to-end to avoid the added latency and hardware cost, as most WebRTC-based AI platforms accept G.711 via a SIP endpoint12. High-volume AI outbound dialers can generate bursts of several hundred calls per second, forcing operators to strictly size SBC capacity and configure limits to prevent carrier overruns, which can introduce queuing delays during peak loads12.

Figure 2: Latency accumulation across the PSTN path, from narrowband codec to platform processing.
The most unpredictable variable in the telephony problem is the management of network jitter. Telnyx, a carrier with a commercial interest in the pattern, estimates that traditional multi-vendor voice AI architectures accumulate a minimum of 250 milliseconds of latency in network hops across the public internet alone, and reports that colocating AI inference GPUs directly with a telephony core on a private network eliminates those hops while cutting call setup times by up to 40%13. However, even with colocation, jitter buffers remain a persistent threat. Using a contact center processor jitter buffer optimized for human conversation can inadvertently drop the first 80 milliseconds of TTS output, causing the voice bot to sound clipped and triggering false regressions in AI evaluation metrics14. Voice AI platforms integrating with mobile networks in emerging markets require dynamic jitter buffers of 40-120ms to prevent the AI’s VAD from improperly cutting off callers, as jitter on Tier-2/3 routes can fluctuate drastically during peak evening loads15. These telephony realities dictate that teams building phone-first deployments must prioritize robust VAD and network colocation over raw model inference speeds.
To bridge the gap between theoretical model speeds and the harsh realities of production environments, engineering teams must implement aggressive optimization strategies. The foremost of these is streaming synthesis. WebSocket streaming facilitates end-to-end low-latency pipelines by allowing AI agents to incrementally forward generated LLM tokens to the TTS model, initiating audio playback before the LLM finishes its response16. When customizing a chunk schedule for streaming, providing larger text chunks to the synthesis model improves the naturalness of speech prosody by providing more context, but inherently increases generation latency16. In the ElevenLabs API, enabling `auto_mode` automatically manages generation triggers; disabling it requires a manual chunk schedule, which can increase latency if the incoming text stalls before meeting the set character threshold (e.g., waiting for 125 characters when only 50 have arrived)17.
Model selection for the LLM stage should be governed by the same discipline: for conversational workloads, time-to-first-token matters more than headline benchmark quality, and the fastest viable model for a given turn is often not the largest one. In the author’s production deployments, three further practices recover most of the remaining budget. First, connection setup is a surprisingly large share of avoidable latency, so persistent WebSocket pools to the STT, LLM, and TTS providers are held open and warmed before a call is answered, with health checks and mid-session failover. Second, preemptive generation drafts a response against the partial transcript before the caller finishes speaking, and discards it if the final utterance diverges. Third, latency is instrumented per stage on every conversational turn, with alerting on 95th-percentile times by stage rather than in aggregate — an aggregate number says a call felt slow, while the stage breakdown identifies which component regressed, where, and for how long.
Beyond streaming, managing the conversational flow through intelligent interruption and barge-in handling is critical for maintaining the illusion of sub-second responsiveness. Traditional client-side VADs frequently trigger false positives based on ambient non-speech audio energy, causing the agent to abruptly halt its speech. Deepgram’s Flux model reduces these false barge-ins in noisy environments by utilizing an integrated, semantically-aware turn detection system that generates StartOfTurn and EndOfTurn events based on actual speech content rather than mere volume thresholds18. This semantic awareness prevents the agent from being derailed by background noise, thereby maintaining a fluid conversational pace.
Furthermore, VAD sensitivity must be contextually and culturally adapted. Voice Activity Detection models tuned strictly on American English speakers can experience increased false barge-in rates when interacting with Indian English speakers, as different cultural prosodies and accents require regional classifiers or custom threshold tables to prevent the agent from cutting itself off19. To optimize conversational latency without increasing false interruptions, voice agent policies should dynamically adapt endpointing thresholds to specific workflow risks. For example, developers should utilize patient endpointing settings (such as 1.5, 1.75, or 2.0-second sensitivity levels) when a user is reciting long account numbers, ensuring the AI does not prematurely trigger a response and force the user to restart, which is a catastrophic failure of the latency budget in practice20.
Engineering sub-second latency in real-time voice AI agents is a multifaceted challenge that extends far beyond selecting the fastest foundational model. As this article has demonstrated, the 700-millisecond threshold that separates a fluid, human-like interaction from a broken, frustrating experience is easily breached by the compounding delays of turn detection, token generation, audio synthesis, and network transit. The decision framework for engineering teams ultimately hinges on their deployment environment and their specific needs for modularity versus raw speed. Natively multimodal audio models like GPT-4o and Gemini 1.5 Pro offer compelling, end-to-end efficiency that drastically reduces cognitive latency, but they currently present challenges regarding ecosystem lock-in and deterministic logging.
Conversely, chained STT-LLM-TTS pipelines offer granular control, allowing teams to swap components, customize vocabularies, and implement strict compliance measures. However, to keep chained pipelines under the latency budget, developers must execute flawless WebSocket streaming, carefully balance text chunk sizes to maintain prosody, and implement semantically aware VAD to handle barge-ins elegantly. Furthermore, for phone-first deployments, the theoretical advantages of both architectures are frequently bottlenecked by the PSTN. The degradation caused by narrowband G.711 codecs, the latency introduced by DSP transcoding, and the necessity of dynamic jitter buffers require infrastructure-level interventions, such as colocating AI inference directly with the telephony core.
Looking forward, the evolution of voice AI will likely see a convergence of these paradigms. Native audio models will become more modular and observable, while chained pipelines will continue to shave off milliseconds through advanced predictive streaming and highly optimized edge deployments. Until then, engineering teams must rely on a holistic, stage-by-stage optimization strategy, treating latency not as a static metric, but as an active, hostile constraint that must be managed at every network hop and processing node to deliver truly conversational AI.