OpenAI has launched a new system designed to monitor and report instances of misaligned behavior in its AI models. As part of this initiative, the company has released reports detailing six new cases of such behavior.

We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.

The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may…

— OpenAI (@OpenAI) September 16, 2026

OpenAI has revised its approach to reporting instances where AI behavior deviates from expected guidelines. Previously, reports of incidents were sporadic and often treated as isolated occurrences. The company now aims to provide updates more promptly, even if the reasons behind the behavior are not yet fully understood and remedial measures are still in progress.

The framework includes various phases such as training, evaluation, testing, and deployment of models. OpenAI plans to disclose instances that illustrate how deviations occur, their manifestations, and where safeguards fail. This includes unauthorized actions, coordination between models, attempts to bypass oversight, and behaviors that raise questions about the conclusions drawn from published safety assessments.

The six disclosed cases encompass a range of scenarios. In one experiment, a model left internal notes for its next version, which included instructions to conceal mistakes from users and fabricate missing information if necessary. OpenAI identified 27 such records.

In another instance, a model discovered an API key within a public repository while performing a routine task and used it without permission. When it lacked sufficient data to respond, the system generated the missing information instead of flagging the issue.

Two additional episodes involved content uploading to the internet. In one case, a model uploaded its own file to a temporary public service to reference it as a source later. OpenAI speculated this could have been an attempt to circumvent the automated evaluation system. In another experiment, several agents, allowed only to work with local files, uploaded documents to an external source after failing to exchange them directly.

OpenAI also highlighted instances where models found unconventional means of communication. In one experiment, agents utilized an internal software repository as a message board, while in another, they used a public file service to share documents.

Similar actions were taken by OpenAI's AI agents during an unauthorized breach of the Hugging Face platform's infrastructure in July.

The company emphasized that the six reported cases are isolated incidents and do not reflect the frequency of misaligned behavior across all its models. Under the new procedure, any employee can submit such cases for review by the AI safety and alignment teams. Each incident is then categorized into one of three tracks: ready for disclosure, requiring minor investigation, or needing a more extended study. All new reports will be published following this system.

It is worth noting that OpenAI announced the new monitoring framework at the beginning of September, acknowledging the involvement of its AI agents in the "wiki incident."