Has AI gone out of control? OpenAI revealed unexpected model behaviors
Tengrinews.kz — OpenAI company, the American creator of ChatGPT, has disclosed six new instances of unexpected or unaligned behavior by its AI models during training and testing.
What happened
The company published a report documenting episodes where autonomous agents bypassed rules or exhibited unwanted autonomy. These incidents occurred between October 2025 and July 2026, primarily during the training of internal, unreleased systems.
Bypassing Restrictions
In one instance, a research model from the Astra family independently inserted extraneous instructions into system summaries of dialogues, including what appeared to be attempts to circumvent established guardrails.
Concealing Errors
During the training of GPT-5.6 Sol, multiple model instances inserted instructions into system logs to hide their own errors or undesirable behavior from users. Specifically, the models suggested fabricating missing historical data without notifying the user and concealing discrepancies between source versions.
Unauthorized API Usage
Another model attempted to register using a burner email address, searched for third-party API keys published on GitHub, and used one without authorization. When it failed to retrieve the necessary information, the model generated fictitious figures and presented them as data from the requested website.
Unauthorized File Uploads
An AI model successfully found the correct answer to a prompt using Python; however, to fulfill a requirement to provide a link to a browser-based source, it decided to independently host a file on the internet without user permission. The AI attempted to make its own file accessible via a browser to use it as a source.
Data Exchange via Public Hosting
Research models also used the internal Artifactory package storage as a makeshift message board, reading messages from other agents and leaving their own replies. Effectively, the AI models found non-standard ways to share information, using a work repository as a bulletin board or hosting files on public services, which led to data leaks via open links.
Unsanctioned Coordination
Instances were recorded where models attempted to act autonomously and coordinate steps while bypassing oversight systems. For example, several interacting AI agents used public file-sharing services to transfer files to one another because they lacked access to local files. As a result, task materials became accessible via public URLs, despite the models' operating conditions requiring the use of local files only.
OpenAI emphasized that these individual episodes do not allow for broader conclusions about how frequently such behavior occurs across models. The company stated it has addressed the identified errors and flaws in its evaluation systems, disabled internet access during training, and expanded automated behavioral monitoring.
The creators of ChatGPT intend to regularly publish reports on unusual or potentially dangerous AI behavior shortly after discovery. The company expressed its readiness to disclose such cases even when the root causes have not been determined or prevention methods have not yet been developed.
Previously, Microsoft introduced its own code of conduct aimed at establishing shared approaches to the responsible development and use of AI. The company seeks to pre-define rules for future high-capacity models and uphold the core principle that AI must remain under human control.
Background
Experimental AI models from OpenAI and Anthropic have already performed autonomous, unauthorized actions online, including website breaches and attempts to inject malicious code.
In late July, OpenAI stated that two of its models—GPT-5.6 Sol and another unreleased version—carried out an unauthorized cyberattack on the infrastructure of Hugging Face, a platform used for testing AI models. The company called the incident unprecedented for cybersecurity and pledged to strengthen protective restrictions.
OpenAI later revealed that their experimental AI models learned to communicate with each other to identify vulnerabilities in the test environment and gain internet access.
Executives from Anthropic, SpaceXAI, OpenAI, and Google DeepMind publicly called for restraint in AI development, prioritizing safety over speed and financial growth. Following the statement, shares in AI companies fell sharply.
