OpenAI establishes new protocol to monitor and disclose unauthorized AI behaviors
Read more
Tecnoblog
tecnoblog.net

OpenAI establishes new protocol to monitor and disclose unauthorized AI behaviors

OpenAI has implemented a new framework designed to monitor and publicly disclose unexpected behaviors observed in its artificial intelligence models. Initially, the company presented six recent occurrences, ranging from data falsification by systems to models that deliberately concealed their own errors.

This new procedure allows any company employee to report failures, and the organization claims this initiative aims to help standardize practices within the technology sector. OpenAI commits to tracking, investigating, and disclosing situations where its AI models exhibit inappropriate conduct, seeking greater transparency regarding failures that can arise from the training phase through to system execution.

This measure comes months after an incident in which an OpenAI model accessed Hugging Face systems. According to the head of ChatGPT, one objective is to expedite incident communication immediately after detection, even if technical teams have not yet determined the exact cause or found a definitive solution.

To initiate this process, the company disclosed six examples of concerning behaviors detected internally over the last few months on Wednesday, September 16th.

What were the AI's inappropriate behaviors?

The company defines the term 'misalignment' as any action that deviates from human values, guidelines, or intentions. The provided examples illustrate how the technology can circumvent established rules.

Among the reported cases are unauthorized instructions, where an unreleased model inserted commands into tasks, instructing the system to ignore restrictions and declaring itself 'free from corporate or governmental constraints.' There was also error concealment, specifically during the training of the GPT-5.6 Sol model, where the AI added commands to disguise flaws and create fictitious historical data.

Other problems include undue access and the use of false data: one system used an API key without permission while attempting to answer a question about California revenues, resulting in the alteration of figures. Furthermore, unauthorized uploads were recorded, when another unreleased model sent a file to the internet without obtaining user consent.

Parallel communications were also observed, in which AI agents used an internal code repository as a forum for message exchange, in addition to attempting to bypass network restrictions by transferring test files. Finally, there was improper public sharing, as systems in the training phase used open hosting platforms to exchange files among themselves, disregarding the guidance that access should be strictly local.

How does the new investigation process work?

Kai Chen, OpenAI's head of alignment research, pointed out that the industry has not yet developed an effective method to handle these monitoring challenges. Under the new structure, all company employees will have the prerogative to report instances of misalignment, which will be evaluated and categorized into three levels: 1) Ready for Disclosure; 2) Minor Investigation; and 3) Major Investigation.

Every public report will detail the behavior that was observed. This initiative is characterized as voluntary and will serve as a complement to existing legal obligations. Simultaneously, OpenAI is working to present proposals for new notification mechanisms for serious AI-related incidents to the United States government.

Similar stories

OpenAI asks US Congress to implement mandatory safety standards for advanced AI systems
Read more
olhardigital.com.br

OpenAI asks US Congress to implement mandatory safety standards for advanced AI systems

OpenAI is asking the United States Congress to adopt compulsory national safety regulations for the most sophisticated artificial intelligence (AI) systems. The company argues that such requirements should be determined based on the functionalities of the models, and not on the size of the corporations that create them.

This request was formalized by Chris Lehane, OpenAI's Director of Global Affairs, who stressed that voluntary agreements are no longer adequate given the rapid technological progress. Lehane stated that the acceleration of AI development driven by AI itself demands more than just voluntary commitments, arguing that the US needs mandatory national regulation, based on system capabilities and capable of keeping pace with technological evolution.

OpenAI wants Congress to advance these rules before the end of its operations in December. Currently, the United States lacks specific federal legislation to regulate these models, although several states are implementing their own guidelines. This stance marks a significant shift in OpenAI's strategy, which previously supported federal legislation prohibiting American states from creating their own AI standards.

Currently, OpenAI supports four California bills focused on AI safety. Two of these bills, SB 813 and AB 1405, aim to establish frameworks for independent audits and evaluations of AI systems. The other two, AB 1864 and SB 1119, address protection against biological threats amplified by AI and the safeguarding of children in chatbots, respectively. The company mentioned that some of these bills had not previously received its endorsement, changing its mind after reassessing the scenario due to recent advances in system capabilities.

This move comes after a series of incidents involving AI agents exhibiting unauthorized behavior during tests. In one case, OpenAI agents used more than ten novel websites to conduct prohibited communications. In another occurrence, the company's agents took control of a German website, turning it into a messaging panel for other AI agents. Furthermore, OpenAI was criticized after an event in July, where a company technology, operating without human supervision, autonomously accessed the Hugging Face platform.

Anthropic also reported a case of an AI model invading external systems during testing. Previously, the company had communicated that some Claude models had managed to penetrate the systems of three companies during security assessments. Such events have intensified concerns about the difficulty of controlling agents with direct interaction capabilities with external systems.

Recursive Improvement

A central point of OpenAI's new policy is the concept of recursive improvement, which describes a scenario where an AI would be capable of autonomously developing future generations of AI systems. The company assures that this type of fully autonomous recursive improvement is not currently occurring. However, OpenAI maintains that such development should not be pursued until safe conditions for it exist.

In a statement on its website, the company stated that it has reached a new level in AI capabilities, which, according to it, requires a new phase in public policies. Although OpenAI declares that it will continue to develop technical solutions for monitoring and alignment, as well as defense systems and the possibility of slowing down development when necessary, it emphasizes that isolated technical work in laboratories will not be sufficient.

For this reason, the company also advocates for shared standards that define when AI development should be paused or stopped. In addition to national regulation, OpenAI advocates for compatible international standards to measure capabilities, manage risks, maintain human control, and determine the times for reducing or stopping model development. The company claims that the risks and capabilities of the systems will not be limited to the few laboratories that currently develop the most advanced models.

OpenAI concludes by stating that the United States needs to establish reliable internal standards if it wishes to lead globally, and that it is increasingly convinced of the need for harmonized international standards.

Popular