Talking and Listening at the Same Time: GPT-Live
OpenAI launches GPT-Live, a full-duplex voice model for ChatGPT delegating complex queries to GPT-5.5, now rolling out globally on iOS, Android and web.
OpenAI is revamping ChatGPT's voice with GPT-Live, a new generation of voice models designed for exchanges that are closer to real conversation. The foundation is a full-duplex architecture: the model listens and speaks simultaneously, inserts attention markers ("uh-huh," "yeah"), responds instantly, or remains silent when the user is thinking.
Two technical approaches underpin the result. First, continuous interaction: instead of processing separate turns, GPT-Live decides multiple times per second whether to speak, listen, interrupt, or call a tool, which also makes live translation possible. Second, delegation: as soon as a question requires a web search or reasoning, the voice model delegates it in the background to a frontier model (GPT-5.5 at launch) and continues the conversation while awaiting the result, with three levels of effort to choose from (Instant, Medium, High).
ChatGPT's voice mode inherits these changes: nine remastered voices, better listening in noisy environments, seamless waiting when hesitating, and visual cards for weather, stock market, or live sports during the exchange. OpenAI claims a clear preference over its previous Advanced Voice Mode in its own tests.
The publisher also details voice-specific security: audio evaluations, dedicated training (self-harm, psychosis, emotional dependence, violence, sexual content), and safeguards capable of acting during speech, even ending the exchange in risky cases, with protections for minors via parental control. The model remains limited to predefined voices, without imitating real people.
The rollout is global on iOS, Android, and ChatGPT.com: GPT-Live-1 becomes the default model for Go, Plus, and Pro offerings, GPT-Live-1 mini for the free offering, with API access announced for soon. At launch, the accent may remain non-native depending on the language, with no video or screen sharing support, and Standard and Advanced Voice modes remaining accessible.