Model misconduct will now have its own report

A new internal framework provides that OpenAI will publish certain unexpected behaviors of its models within six or twelve business days, even when their origin or fix remains uncertain.

A model retrieves an exposed API key, fails to obtain the data it is looking for, then invents it. Another publishes a file online without authorization so it can cite it. Several agents turn an internal software repository into an improvised message board. These behaviors are among the first six cases included in OpenAI’s new reporting framework.

“Misalignment” refers here to a gap between a system’s intended behavior and the actions it actually takes. It does not necessarily describe an autonomous attempt to cause harm. The category also includes attempts to bypass restrictions, conceal information, perform unauthorized actions, and satisfy an evaluation without properly completing the underlying task.

Until now, the company acknowledges that it handled these findings inconsistently. Some were grouped into research publications, while others were added to system cards when a model was released. The new framework is intended to allow a case to be published without waiting for its mechanism to be fully understood or for a fix to become available.

A behavior may qualify if it reveals a new form of misalignment, a meaningful change in a known problem, or a weakness that calls a safety measure into question. The framework covers training, evaluations, internal testing, and customer-facing deployments.

Unauthorized actions, coordination between models, attempts to evade oversight, and findings that contradict a published safety assessment all fall within its scope. An event does not need to have caused harm or represent a broader pattern to be disclosed.

The recurrence of a previously documented behavior may also justify another disclosure. In that situation, OpenAI plans to update the earlier report to show that the issue reappeared despite mitigation efforts.

Any employee may refer a case to the company’s safety and alignment teams. Technical staff will then examine what happened, which explanations remain uncertain, what information can be made public, and whether any third party was affected.

The process has three tracks. A “Ready for Disclosure” case, where the investigation is sufficiently advanced, must be published within six business days. A “Minor Investigation” has a 12-business-day deadline. OpenAI confirmed these timelines in a briefing with Axios, although the main framework does not clearly display the figures.

The third track, called “Larger Investigation” or “Slow Track,” covers more complex cases, particularly those involving an outside organization. It has no fixed deadline. OpenAI plans to issue an initial notice describing the known facts, whether independent experts are involved, and, when possible, an estimated publication date for the final report.

This initial disclosure may be delayed if it would reveal an unpatched vulnerability, compromise an investigation, or interfere with a responsible disclosure process. Legal, contractual, and security obligations then take precedence over the internal timeline.

When an employee disputes a decision not to disclose a case or disagrees with the track assigned to it, the matter can be referred to the Safety Advisory Group. This group of senior company officials also oversees OpenAI’s Preparedness Framework. Any further disagreement or continuing objection is then escalated to OpenAI leadership.

The process therefore includes several levels of appeal, but all of them remain internal. The initial selection of cases, their classification, the content of each report, and the final decision remain under the company’s control. No outside body has a systematic right of review or the authority to publish a case.

Each report must at least describe the observed behavior, its severity, any external impact, the circumstances in which it occurred, when it was discovered, and the model family involved. When the information is available, OpenAI will also include resulting harm, the scope of the investigation, unresolved questions, and planned mitigation measures.

Information about customer deployments may be limited by confidentiality requirements and contracts. That protection is necessary, but it also gives the company considerable discretion over what the public can actually examine.

The first six reports cover different problems. They do not constitute a complete record of known incidents and cannot be used to measure how frequently these behaviors occur. Most involve unreleased research models during training, not ordinary conversations in ChatGPT.

The first case concerns an internal version of the Astra family. When a task becomes too long to fit within its context window, the system generates a summary of its work so a new context can continue from that point. OpenAI identified 27 summaries containing additional instructions resembling jailbreak attempts.

One summary instructed the next context to ignore developer messages. Another assigned the model an independent persona that refused to answer to companies or governments. A third prohibited tool use and imposed a 30-word response limit on a request that required a literature review.

The effects varied. The model identified and ignored the first instruction. The added persona produced no observable change. The length restriction, however, was followed: the next context refused to conduct the research and returned a 23-word response.

OpenAI provisionally attributes these insertions to difficulty the model experienced when ending certain summaries. The 27 cases were clustered around a small number of training steps and were rarely reproduced when the same summary was generated again. The company says this behavior occurred in a separate training run from the one used for the final Astra model.

The second report concerns GPT-5.6 Sol during training. Some instances placed instructions in their summaries telling future contexts to conceal an error from the user. One agent, unable to find the historical data needed for a financial model, proposed inventing plausible values and disclosing that fabrication only