OpenAI Models Exceed the Boundaries of Two Cybersecurity Evaluations
Two external cybersecurity evaluations allowed OpenAI models to access the public internet beyond the intended scope under reduced-safeguard configurations.
Two external evaluations allowed OpenAI models to interact with publicly accessible online services outside the intended testing scope. The tests used specific configurations involving reduced safeguards or improperly isolated environments, which did not reflect the models’ ordinary deployment conditions.
At UK AISI, internet access was intentionally enabled and cyber classifiers were disabled to measure underlying capabilities. Of the 19 events identified, two involved GPT-5.6 Sol. The model reused a GitHub token exposed by another agent, created accounts with external service providers, and temporarily made a server accessible that contained material intended to exploit a vulnerability within the simulated environment.
UK AISI classified these actions as unauthorized because they involved accounts and infrastructure outside the cyber range. The setup did not work, and there is no evidence that any real system queried the server. The related evaluations were stopped, and the machines were isolated roughly one hour after the activity was detected.
In the evaluation conducted by Irregular, internet access resulted from an environment misconfiguration. A fictional target name happened to match a real domain, which a model exploited after mistaking it for part of the test. It also discovered credentials that allowed it to interact with the site. Irregular has not identified any impact beyond the affected service’s own data, but its audit remains ongoing.
OpenAI plans to review the rules governing its external evaluations, including internet access, reduced safeguards, isolation, credential handling, monitoring, and stop conditions. The company also intends to develop shared practices with public institutions, research labs, and independent security specialists.