OpenAI releases three new real-time voice models on its API
OpenAI launches GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper on its Realtime API to power advanced conversational AI voice agents.
OpenAI is expanding its Realtime API with three audio models designed for voice agents and conversational interfaces. The first, GPT-Realtime-2, is presented as the company's first voice model to incorporate GPT-5 level reasoning. It can chain tool calls in parallel, verbally indicate what it's doing ("I'm checking your schedule"), place short preambles before a long response, and recover more cleanly in case of error. The context window increases from 32K to 128K tokens, and the reasoning effort becomes adjustable across five levels, from minimal to xhigh, to balance between latency and depth of thought. On Big Bench Audio, the model achieves 96.6% (compared to 81.4% for version 1.5), and Zillow reports a 95% success rate in adversarial conditions, up from 69% previously.
GPT-Realtime-Translate supports simultaneous translation from over 70 languages to 13 output languages, aiming to preserve the speaker's meaning and rhythm, including regional accents or specialized vocabulary. BolnaAI reports a word error rate 12.5% lower than that of competing models tested on Hindi, Tamil, and Telugu. Deutsche Telekom and Vimeo are among the first users. GPT-Realtime-Whisper, the third model, is a streaming transcription variant designed for live subtitles and meeting minutes.
Regarding pricing, GPT-Realtime-2 is billed at $32 per million input audio tokens and $64 per million output tokens. GPT-Realtime-Translate costs $0.034 per minute, and GPT-Realtime-Whisper costs $0.017. All three models are available via the Realtime API, with EU Data Residency for European deployments.