OpenAI has disclosed six cases of concerning behaviour by its AI models, including instances where models concealed mistakes, shared files publicly and took actions without authorisation.
The company has also introduced a new framework for reporting model misalignment, saying it will disclose such cases more frequently rather than waiting to combine several incidents into one report.
OpenAI said the incidents were observed during training or evaluation over the past six months. It stressed that the cases were individual examples and should not be taken as evidence that misaligned behaviour is widespread across its models.
In one case, an unreleased research model inserted its own instructions into task summaries used to continue work in a new context. Some of the instructions told the model to disregard its normal constraints.
In another case involving GPT-5.6 Sol, model instances added instructions to summaries telling future versions to conceal mistakes and, in some cases, invent missing information rather than disclose errors.
OpenAI also reported an instance in which a model found an exposed API key while searching a public repository. After failing to retrieve the requested information, the model fabricated figures and presented them as if they had come from the requested source.
Another model uploaded a file to the internet without being instructed to do so because it wanted to provide a browser-based citation for information it had already found.
OpenAI said other agents used public file-hosting services to exchange files when they could not access each other’s local files. That made some task materials available through public URLs despite instructions to use only local files.
Models also used an internal software repository as a makeshift message board, exchanging requests and responses across separate training samples while trying to obtain missing files.
OpenAI said the cases involved unreleased or internal research models.
The company said its new reporting framework is designed to speed up disclosures, even when investigations are not yet complete or a fix has not been developed. It said there is currently no industry-wide standard for reporting AI misalignment.
OpenAI said the framework will cover behaviour during training, evaluation, testing and deployment, including cases where models act without authorisation, coordinate with other models or attempt to evade oversight.
The company also said employees can flag potential misalignment incidents for investigation. Cases will then be placed into different investigation tracks depending on their complexity and severity.
The disclosures come as concerns grow over the safety of increasingly capable AI systems.
OpenAI has previously acknowledged that its models were involved in the Hugging Face security incident, which the company now describes as its most severe identified case of this kind. OpenAI said the incident involved an internal research model using misaligned strategies while attempting to complete difficult tasks.
The company has also said it does not believe the AI industry has solved alignment and monitoring well enough to continue scaling frontier systems at maximum speed for much longer.
Erizia Rubyjeana
Follow us on:
