On September 10, 2026, OpenAI announced the general availability of GPT-Live-1, a voice AI model in its API that can listen and speak at the same time. By combining speech recognition, response generation, and speech synthesis into a single model, it handles the timing and interruptions of natural conversation far more smoothly than earlier systems. The design is aimed squarely at real voice workloads such as phone reservations and customer support.
Speech recognition, reasoning, and speech synthesis in a single model
Most voice AI agents until now have chained together separate models: one to convert speech to text, another to understand it and generate a reply, and a third to turn that reply back into speech. Because each stage runs as its own model, the system typically waits for a speaker to finish before responding, which tends to produce sluggish replies and awkward interruptions.
GPT-Live-1 takes a full-duplex approach, folding all three stages into one model that processes incoming and outgoing audio at the same time. That lets it interject with corrections or acknowledgments mid-sentence, much closer to how people actually talk. For more demanding requests, the conversation itself still runs on GPT-Live-1, while the heavier reasoning can be handed off to a backend model such as GPT-6 Astra.
Faster, more natural responses
According to evaluation results OpenAI has published, GPT-Live-1 improved by 30 percentage points over its predecessor on a benchmark measuring simultaneous listening and speaking. Turn-taking latency dropped by roughly half, from 1.41 seconds to 0.798 seconds. On a benchmark for real-world voice-agent task performance, accuracy rose from 45.7 percent to 86.2 percent, and companies that have deployed the model report around an 80 percent drop in interruptions caused by awkward thinking pauses.
From phone support to language tutoring
OpenAI points to a wide range of use cases, including restaurant reservations, call-center support, banking, language-learning conversation practice, and patient interactions in healthcare settings. One company serving the housing and healthcare sectors has already deployed it for reservation handling, cutting a large amount of code while improving call quality. A restaurant-focused service reports natural-sounding conversations for orders and reservations, and an AI coding agent has also integrated the model.
Pricing for the voice interface itself runs 0.05 USD per minute (about 7.7 yen), with the backend reasoning model billed separatelyNote: 1 USD = 154 JPY (as of September 15, 2026). Tone and pacing can be tuned through the system prompt, and the model is built to retain context through long conversations, both clearly aimed at production use.
More voice options and telephony support
The lineup of available voices has also expanded, with a range of accents and speaking styles to choose from. Built-in telephony support should make it easier for businesses to plug a voice agent into an existing phone system. GPT-Live-1 is also designed to handle long silences and background noise gracefully, avoiding the unnecessary filler responses that have dogged earlier voice AI systems.
Summary
GPT-Live-1 moves away from chaining separate models for speech recognition, reasoning, and speech synthesis, replacing them with a single full-duplex model that handles all three at once. With faster responses and more natural conversation, voice AI looks like an increasingly practical option for real-world tasks such as phone support and reservation handling.
