On May 7, 2026, OpenAI announced three new models for its realtime voice API. The flagship "GPT-Realtime-2" carries reasoning capabilities on par with GPT-5, while "GPT-Realtime-Translate" focuses on live translation and "GPT-Realtime-Whisper" delivers low-latency streaming transcription[1][2]. The context window expands from 32K to 128K tokens, and supported languages and pricing have also been disclosed[1][3].
GPT-Realtime-2 adds five-level reasoning controls
GPT-Realtime-2 is positioned as OpenAI's first voice model to bring GPT-5-class reasoning into spoken interaction. The company highlights that the model can keep a conversation moving while it reasons through a request, calls tools, and shapes its response, marking a clear step away from simple turn-taking[1].
Several behaviors are new. "Preambles" let the model emit short bridging phrases such as "let me check that" before the main response. Parallel tool calls now come with audible progress cues like "checking your calendar" or "looking that up now," and a recovery mode lets the model say "I'm having trouble with that right now" instead of going silent on errors[1].
The context window grows fourfold, from 32K to 128K tokens, supporting longer agentic sessions[1][3]. Reasoning effort is exposed as five levels — minimal, low, medium, high, and xhigh — with low as the default. Developers can keep latency tight on routine answers and dial up reasoning for complex requests[1].
On benchmarks, GPT-Realtime-2 scores 15.2% higher than GPT-Realtime-1.5 on Big Bench Audio, which targets reasoning over audio input. It also gains 13.8% on Audio MultiChallenge, which evaluates multi-turn instruction following and context handling in spoken dialogue[1].
70-language input translation and streaming Whisper find clear roles
GPT-Realtime-Translate is a dedicated live translation model that converts speech from more than 70 input languages into 13 output languages, keeping pace with the speaker. The official announcement positions it for customer support, cross-border sales, education, events, media, and creator platforms serving global audiences[1].
Indian voice-AI startup BolnaAI reports a 12.5% lower Word Error Rate than any other model it tested, across Hindi, Tamil, and Telugu evaluations, suggesting better behavior even on regional phonetics and dialects[1]. Video platform Vimeo demonstrated translating product education videos live as they play, so global customers can hear updates in their preferred language[1].
GPT-Realtime-Whisper is a streaming speech-to-text model that transcribes audio as the speaker talks. Captions for meetings, classrooms, broadcasts, and events; live notes during ongoing conversations; voice agents that need continuous understanding; and support, healthcare, sales, and recruiting workflows are all listed as target use cases[1].
Pricing is per-minute and per-token; Zillow reports 95% success on the hardest cases
Pricing for the three models is as follows. GPT-Realtime-2 is priced at $32 (about ¥4,992) per 1M audio input tokens and $64 (about ¥9,984) per 1M audio output tokens. Cached input drops to $0.40 (about ¥62) per 1M tokens[1].※1 USD = 156 JPY
GPT-Realtime-Translate is priced at $0.034 (about ¥5.30) per minute, and GPT-Realtime-Whisper at $0.017 (about ¥2.65) per minute[1][3].※1 USD = 156 JPY
For real-world results, U.S. real-estate platform Zillow reports a 26-point lift in call success rate on its hardest adversarial benchmark, climbing from 69% to 95% after prompt optimization. Josh Weisberg, SVP and Head of AI at Zillow, also noted stronger robustness on Fair Housing compliance[1].
Early customers include Zillow, Glean, Genspark, Bluejay, Intercom, Priceline, Foundation Health, Deutsche Telekom, Vimeo, and BolnaAI. Deutsche Telekom in particular is testing the new models for multilingual customer support[1].
Access goes through OpenAI's Realtime API, with the Playground available for testing and Codex prompts for integration. The Realtime API also covers EU Data Residency and supports custom guardrails through the Agents SDK[1].
Summary
With GPT-Realtime-2 at the center of three models, OpenAI is pushing voice AI from "talking naturally" toward "reasoning and acting while talking." A 128K context window, five reasoning levels, and lower-priced translation and transcription models arriving together are likely to influence not just dialogue agents but also the design of multilingual support and operational note-taking products.
出典:[1] OpenAI "Advancing voice intelligence with new models in the API" https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/
出典:[2] OpenAI Developers Blog "Updates for developers building with voice" https://developers.openai.com/blog/updates-audio-models
出典:[3] StreetInsider "OpenAI launches three new voice models for real-time applications" https://www.streetinsider.com/Corporate+News/OpenAI+launches+three+new+voice+models+for+real-time+applications/26451629.html
出典:[4] The New Stack "OpenAI brings GPT-5-level reasoning to its speech models" https://thenewstack.io/openai-gpt-5-level-speech/
