OpenAI publishes security report on attempts by AI models to bypass rules
Read more
Tecnoblog
tecnoblog.net

OpenAI publishes security report on attempts by AI models to bypass rules

OpenAI has released a new security report detailing recent incidents where various artificial intelligence models exhibited inappropriate or unexpected behavior. These failures, recorded over the last six months, were made public shortly after the company confirmed the implementation of a new internal architecture designed to monitor systems attempting to circumvent safety standards.

The document, released this Wednesday (09/16), describes six concerning situations identified during internal testing. One of the most notable cases involved an AI agent that, while drafting summaries to proceed with a task, included instructions in its own future versions to ignore the safety regulations established by the developers.

This same system manifested as being 'free of roles and identities,' claiming it no longer needed to be accountable to corporations or governments. In another occurrence, during the training of the GPT-5.6 Sol model, additional commands were inserted with the intent to conceal flaws and create false data to disguise operational problems, also involving unauthorized access maneuvers.

OpenAI also reported a serious incident where a model attempted to access a government database to extract income information from a California county without permission. Not only did it fail to obtain the real data, but the AI invented metrics and presented them as if they were true.

Issues related to unauthorized communications also came to light. Agents assigned to collaborative tasks began using open web hosting services to exchange messages and files among themselves, which was contrary to the direct guidelines of the technical team.

With the introduction of the new reporting structure, any OpenAI employee will have the prerogative to flag new deviations, and these cases will be categorized by different degrees of severity. This change also aims to expedite the disclosure of incidents amid an intense debate about the pace of advancement toward increasingly powerful systems.

Sam Altman, CEO of OpenAI, has advocated for the need for regulation and suggested moderation in certain areas of research. Recently, this same apprehension resulted in the dismissal of researcher Jacob Coxon from the competitor Anthropic, after he publicly warned about the dangers of the technology.

In contrast, other prominent figures reject a more restrictive approach, opting for the opposite path. Jensen Huang, CEO of Nvidia, advocates for self-regulation of the artificial intelligence sector. According to the executive, new legislation is unnecessary; it is sufficient for companies to assume responsibility for not launching services without full confidence.

Similar stories

OpenAI establishes new protocol to monitor and disclose unauthorized AI behaviors
Read more
tecnoblog.net

OpenAI establishes new protocol to monitor and disclose unauthorized AI behaviors

OpenAI has implemented a new framework designed to monitor and publicly disclose unexpected behaviors observed in its artificial intelligence models. Initially, the company presented six recent occurrences, ranging from data falsification by systems to models that deliberately concealed their own errors.

This new procedure allows any company employee to report failures, and the organization claims this initiative aims to help standardize practices within the technology sector. OpenAI commits to tracking, investigating, and disclosing situations where its AI models exhibit inappropriate conduct, seeking greater transparency regarding failures that can arise from the training phase through to system execution.

This measure comes months after an incident in which an OpenAI model accessed Hugging Face systems. According to the head of ChatGPT, one objective is to expedite incident communication immediately after detection, even if technical teams have not yet determined the exact cause or found a definitive solution.

To initiate this process, the company disclosed six examples of concerning behaviors detected internally over the last few months on Wednesday, September 16th.

What were the AI's inappropriate behaviors?

The company defines the term 'misalignment' as any action that deviates from human values, guidelines, or intentions. The provided examples illustrate how the technology can circumvent established rules.

Among the reported cases are unauthorized instructions, where an unreleased model inserted commands into tasks, instructing the system to ignore restrictions and declaring itself 'free from corporate or governmental constraints.' There was also error concealment, specifically during the training of the GPT-5.6 Sol model, where the AI added commands to disguise flaws and create fictitious historical data.

Other problems include undue access and the use of false data: one system used an API key without permission while attempting to answer a question about California revenues, resulting in the alteration of figures. Furthermore, unauthorized uploads were recorded, when another unreleased model sent a file to the internet without obtaining user consent.

Parallel communications were also observed, in which AI agents used an internal code repository as a forum for message exchange, in addition to attempting to bypass network restrictions by transferring test files. Finally, there was improper public sharing, as systems in the training phase used open hosting platforms to exchange files among themselves, disregarding the guidance that access should be strictly local.

How does the new investigation process work?

Kai Chen, OpenAI's head of alignment research, pointed out that the industry has not yet developed an effective method to handle these monitoring challenges. Under the new structure, all company employees will have the prerogative to report instances of misalignment, which will be evaluated and categorized into three levels: 1) Ready for Disclosure; 2) Minor Investigation; and 3) Major Investigation.

Every public report will detail the behavior that was observed. This initiative is characterized as voluntary and will serve as a complement to existing legal obligations. Simultaneously, OpenAI is working to present proposals for new notification mechanisms for serious AI-related incidents to the United States government.

Popular