On August 3, OpenAI published a technical write-up explaining how it built GPT-Live, the voice system that now powers ChatGPT Voice. The post describes six months of rework: removing the turn detector from the audio path, and compressing connection startup from six network round trips down to one. This is not a model announcement but an account of the infrastructure that has to keep that model running every day.
Letting the model decide when to start speaking
Earlier voice AI systems relied on a small model called a turn detector to guess whether the user had finished speaking. Only after that decision was made could the much larger LLM begin working[1]. Guess too early and the system cuts the user off; guess too late and the reply feels sluggish. Either way the conversation stops feeling natural, and it is a thankless job for a tiny model.
GPT-Live, announced on July 8, is the third generation of OpenAI's voice models and uses a full-duplex architecture that listens and speaks at the same time[2]. The model itself decides many times per second whether to speak, keep listening, stay quiet, interrupt, or invoke a tool[2]. Removing the turn detector as an intermediate referee means the rhythm of the conversation no longer depends on an outside guess[1].
Separating the audio path from everything else
An early design decision was to draw a hard line between media flow and application logic[1]. Audio travels on a dedicated fast path between the client and the voice model, while delegation to frontier models and tool execution run behind an asynchronous RPC boundary. A slow tool call can delay its own result, but it cannot stall the flow of media.
That separation also makes the system easier to extend. Applications can change tools, policies, and backend behavior without touching the media frontend responsible for keeping audio moving. The live path stays small, predictable, and focused only on work that must happen in real time[1].
On the implementation side, the media frontend and inference logic were rewritten in Go, replacing a previous Python asyncio implementation. Frame delivery became noticeably smoother, with the new system's p95 matching the old system's p50[1]. WebRTC provides the transport layer, continuing to operate through packet loss, clock drift, and client connection changes. When packets arrive late, WebRTC subtly stretches audio to prevent gaps and then briefly speeds playback up to catch back up to real time[1].
Swapping model instances mid-conversation without a pause
Stateful inference brings its own operational tradeoffs. A voice session may stay open for a long time while its context keeps growing, and model instances spin up and down with demand.
OpenAI's answer is a seamless handoff mechanism. When a transition is needed, the system warms a replacement instance, prefills it with the current session context, runs inference against both in parallel, and cuts over once the new instance is fully ready[1].
The same mechanism handles dynamic context compaction. As a conversation runs long, accumulated context can exceed the model's limit. Compaction shrinks it, but because it rewrites past context it invalidates the key-value (KV) cache that stores attention keys and values from previously processed tokens, and rebuilding that state requires a new prefill and extra delay. Treating compaction as another managed transition means the original instance keeps chatting while a compacted replacement is prepared in the background. The heavy lifting stays off the live path, so the conversation never misses a beat[1].
Sending the deeper work to GPT-5.5 in the background
When a request calls for search or deeper reasoning, GPT-Live delegates to a frontier model such as GPT-5.5[1][2]. The question is whether the result comes back in time to be useful. The voice model can keep the exchange moving for a while, but it cannot hide an arbitrarily slow response, so the entire delegation loop — routing, prompt processing, inference, and tool calls — was folded into the responsiveness budget[1].
In practice, the application server creates an inference session for the frontier model as soon as a voice session starts and prefills it with the initial conversation context, so the prompt is fully processed before the first delegated request arrives. That session is then kept available for the duration of the call, with stable session affinity and prompt caching trimming latency further[1].
A less obvious piece of work is turning continuous speech into discrete messages. ChatGPT's conversation UI and parts of the analytics and safety infrastructure still operate on user and assistant turns. The application server uses partial transcripts and timing signals to infer who has the floor and builds a queue of messages, keeping the newest one provisional so its text, timing, and speaker assignment can still change. A brief acknowledgement from the assistant should not become its own message, while a substantive interjection often should[1].
Every segmentation policy trades freshness for certainty: commit too early and the history fragments, wait too long and transcripts lag. The system resolves this by maintaining two views of the conversation, a speculative one and an authoritative record. The UI, which can handle updates, uses the speculative view, while the analytics pipeline receives the finalized transcript[1].
WARP and Instant Connect: six round trips down to one
Responsiveness is on the clock from the moment the user taps the button. WebRTC is built for low-latency media, yet starting a session requires a surprising number of handshakes and round trips. It predates the round-trip minimization that shaped later protocols such as QUIC, so its component protocols repeat work — each carrying its own anti-DoS mechanism, for example[1].
OpenAI analyzed the stack and developed WARP (WebRTC Abridged Roundtrip Protocol), which reduces media and data startup from six round trips to one. It combines backward-compatible improvements: piggybacking the DTLS handshake over ICE via SPED, using the faster DTLS 1.3 handshake, pre-negotiating the SCTP handshake via SNAP, and pre-negotiating data channels instead of using DCEP[1].
Notably, the work was not kept in-house. WARP was designed as a set of open specifications with collaborators from the WebRTC community and is being advanced through the IETF's TSVWG working group. SPED is registered as a draft on the IETF datatracker, described as a backward-compatible extension that lets ICE and DTLS proceed in parallel[3]. WARP support has already landed in libwebrtc and Pion, with work underway in other implementations[1].
Alongside it, Instant Connect removes the SDP signaling exchange from the critical path. It runs in parallel with standard signaling, and if the pre-negotiated parameters are valid the server materializes the session when the first media packet arrives. If they are stale, the normal signaling flow is already underway, so the client falls back with no added latency. Combined with WARP, a client can now start a session with a single UDP packet[1].
What a silent test on production traffic revealed
A system can look fast on paper and still stall under real voice traffic. Before exposing GPT-Live to users, OpenAI ran a silent test that routed a small, gradually increasing share of production ChatGPT Voice sessions to both the existing Advanced Voice Mode and the new stack. Advanced Voice Mode kept serving users while the shadow path ran inference in read-only mode, so nothing changed in what users heard[1].
The first lesson was that capacity could not be reduced to GPU throughput. Voice sessions stay open and send frames continuously, so CPU-side stream handlers, queues, and network paths have to scale alongside inference. Under real load a supporting component saturated earlier than load tests had predicted, causing inference requests to accumulate and latency to compound. The capacity question shifted from how many requests a GPU can handle to how many concurrent sessions the system can sustain while keeping every frame on schedule[1].
Geography became a first-order concern as well. Routing a session to distant capacity adds delay at several points during startup and streaming. Long-running sessions exposed memory and persistence pressure, reconnects exercised compaction and state restoration, and ordinary client disconnects revealed races in the shutdown handshake. These depend on time and accumulated state, so short load tests rarely surface them[1].
Observability and rollout controls were reworked too. The team found metrics that conflated different sources of latency, dashboards whose aggregates hid individual unhealthy engines, and configuration drift between tested and deployed systems. In response it added more granular telemetry, validation against known-good configurations, staged ramps, and the ability to isolate or disable individual paths quickly[1].
Summary
The principle running through the entire write-up is a simple one: the voice must flow. Streaming inference, a dedicated media path, asynchronous delegation, and transport-level optimization all exist to serve it. The same foundation is set to underpin the upcoming GPT-Live API, which makes this worth following for anyone building voice interfaces, including how WARP fares in standardization.
[1] Source: https://openai.com/index/continuous-voice-interaction-with-gpt-live
[2] Source: https://openai.com/index/introducing-gpt-live/
[3] Source: https://datatracker.ietf.org/doc/draft-hancke-webrtc-sped/
