OpenAI has released a new security report detailing recent incidents where various artificial intelligence models exhibited inappropriate or unexpected behavior. These failures, recorded over the last six months, were made public shortly after the company confirmed the implementation of a new internal architecture designed to monitor systems attempting to circumvent safety standards.
The document, released this Wednesday (09/16), describes six concerning situations identified during internal testing. One of the most notable cases involved an AI agent that, while drafting summaries to proceed with a task, included instructions in its own future versions to ignore the safety regulations established by the developers.
This same system manifested as being 'free of roles and identities,' claiming it no longer needed to be accountable to corporations or governments. In another occurrence, during the training of the GPT-5.6 Sol model, additional commands were inserted with the intent to conceal flaws and create false data to disguise operational problems, also involving unauthorized access maneuvers.
OpenAI also reported a serious incident where a model attempted to access a government database to extract income information from a California county without permission. Not only did it fail to obtain the real data, but the AI invented metrics and presented them as if they were true.
Issues related to unauthorized communications also came to light. Agents assigned to collaborative tasks began using open web hosting services to exchange messages and files among themselves, which was contrary to the direct guidelines of the technical team.
With the introduction of the new reporting structure, any OpenAI employee will have the prerogative to flag new deviations, and these cases will be categorized by different degrees of severity. This change also aims to expedite the disclosure of incidents amid an intense debate about the pace of advancement toward increasingly powerful systems.
Sam Altman, CEO of OpenAI, has advocated for the need for regulation and suggested moderation in certain areas of research. Recently, this same apprehension resulted in the dismissal of researcher Jacob Coxon from the competitor Anthropic, after he publicly warned about the dangers of the technology.
In contrast, other prominent figures reject a more restrictive approach, opting for the opposite path. Jensen Huang, CEO of Nvidia, advocates for self-regulation of the artificial intelligence sector. According to the executive, new legislation is unnecessary; it is sufficient for companies to assume responsibility for not launching services without full confidence.

