GPT-Live offloads heavy tasks and inherits the security of the models it calls.
OpenAI's GPT-Live-1 and GPT-Live-1 mini voice models delegate complex tasks to other models to inherit their safety training, according to a new system card.
OpenAI is accompanying its voice models, GPT-Live-1 and GPT-Live-1 mini, with a system card focused on the risks specific to voice.
A key architectural point structures the document: these voice models delegate complex tasks to other in-house models, and the resulting work then inherits the safety training of the solicited model. The published evaluations therefore describe behavior with delegation, to align with the real deployment context. In addition to the safety already in place for text, safeguards designed for oral interaction are added: controlled inputs and outputs throughout the exchange, the ability to influence or interrupt a response, to broadcast a vocal safety message, to display written resources, or, in the most high-risk cases, to terminate the call.
To measure all of this, OpenAI states it has built voice-specific evaluations, based on real user cases who agreed to share their exchanges and on synthetic prompts targeting edge situations. The company specifies that these datasets are intentionally difficult and adversarial, not representative of rates observed in common use. On these tests, GPT-Live performs as well as or better than the previous Advanced Voice Mode, with two slight regressions (emotional dependency, sexual content) presented as non-significant.
Regarding the Preparedness Framework, the Safety Advisory Group concludes that neither model, without delegation, can be judged "High" in the monitored categories (biological and chemical risk, AI self-improvement, cybersecurity). Cyber exposure remains limited by the absence of code execution and autonomous access to tools. Red teaming, conducted internally and externally without system protections, covered identity impersonation, self-harm, emotional dependency, and child-like voices.