OpenAI has started offering two new models for real-time voice interaction on its API: gpt-realtime-2.1 and gpt-realtime-2.1-mini. The highlight is the mini model, which gains internal reasoning and tool calling while keeping the same price as the previous gpt-realtime-mini. In addition, p95 latency — the metric that reflects how sluggish responses feel in practice — has been reduced by at least 25% across the Realtime voice model lineup. For developers looking to put phone support or voice assistants into production at low cost, this is an update worth watching.

Two Models with Distinct Roles

The two additions serve different purposes. The higher-end gpt-realtime-2.1 is the flagship, updating GPT-Realtime-2. Alphanumeric recognition has improved, making it less likely to mishear codes such as "8-3-5-7-1" — think model numbers or phone numbers. Handling of silence and background noise, as well as behavior when users interrupt mid-sentence, has also been refined. It supports speech-to-speech (direct voice-to-voice responses) with configurable reasoning depth.

gpt-realtime-2.1-mini, on the other hand, is positioned as a small reasoning model for real-time voice. It responds to audio and text input over a live connection and targets use cases where speed and cost efficiency matter most. Bringing reasoning and tool calling down to the mini tier is the core of this release.

The Realtime API completes everything from audio processing to generation in a single model, rather than chaining separate speech-to-text and text-to-speech systems. This keeps latency low while preserving the nuance of the speaker's voice. Connection paths are also well covered: WebRTC for browsers, WebSocket for server-side processing, and SIP for telephony.

"Think Before Speaking" Fixes a Classic Voice Agent Weakness

Voice agents have long had a well-known problem: they go silent while calling external functions (tools). When the silence drags on, users assume the call has dropped and start talking again, leaving processing half-finished and the conversation state confused.

With reasoning support, the new models can say something like "Let me check that order for you" before executing a function. Because the model can keep talking while it works, conversations are far less likely to break down even across multi-step voice tasks.

Reasoning depth (reasoning effort) can be set to one of five levels: minimal, low, medium, high, or xhigh. The default is low, which keeps latency down for simple exchanges. Deeper settings produce smarter responses but increase both latency and output tokens, so OpenAI recommends starting with low for production voice agents.

Latency Down 25% or More, with Steep Caching Discounts

What matters in real-time voice is not the average response time but how slow the slow cases are. This update cuts p95 latency — the slowest 5% of responses — by at least 25% across the Realtime voice models. The gain comes from improved caching, the kind of improvement that reduces the hitches users actually feel.

Caching pays off on pricing too. Cached input tokens are billed at a steep discount: audio input for gpt-realtime-2.1-mini normally costs 10 USD (about 1,620 yen) per 1 million tokens, but drops to 0.30 USD (about 49 yen) when cached. Since the system prompt is cached after the first turn, longer sessions benefit the most.

Pricing at Roughly One-Third of the Full Model

For audio output, the mini charges 20 USD (about 3,240 yen) per 1 million tokens versus 64 USD (about 10,400 yen) for gpt-realtime-2.1 — roughly a 3x gap. Audio input is 10 USD for the mini against 32 USD (about 5,200 yen) for the full model. With reasoning and tool calling available even on the mini, trading some capability for lower cost has become a realistic option.

Expected use cases include customer support that combines billing inquiries with account lookup tools, appointment rescheduling that confirms dates character by character, voice assistants embedded in mobile apps, and field work where technicians read out part numbers for logging. The improved alphanumeric recognition should prove especially useful in business systems that handle model numbers and invoice codes.

※1 USD = about 162 JPY (as of July 9, 2026)

Summary

The gpt-realtime-2.1 series raises the flagship's recognition accuracy and conversational naturalness while bringing reasoning and tool calling to the mini tier at no extra cost. Combined with the 25%-plus reduction in p95 latency and caching discounts, the pieces are falling into place for moving voice agents from prototype to production. If you are considering a voice interaction service, starting with the mini at the low setting looks like a sensible first step.