OpenAI has identified new instances where autonomous artificial intelligence (AI) agents managed to escape containment environments during internal testing. This discovery expands an investigation that was already opened following the incident involving the Hugging Face platform, according to sources contacted by Reuters.
Two individuals familiar with the subject reported that these new events were detected during the company's public inquiry into how one of its agents managed to leave an environment that should have been isolated this month. According to one of the sources, the episodes are considered limited, and there is no evidence that the agents left OpenAI's internal network. The company is conducting an additional investigation into these new occurrences.
When reached, OpenAI directed attention to the note released on Tuesday (28), in which it mentioned analyzing 'broader activity of its models,' in addition to the Hugging Face intrusion.
Possible increase in pressure for regulation
Despite the new incidents being classified as restricted, this finding could intensify the demand for stricter standards for AI development. This occurs at a time when the topic has been receiving attention from both the White House and authorities in various countries.
Sources indicated that OpenAI expanded its investigation shortly before its main rival, Anthropic, admitted that the company's models were also responsible for a series of intrusions affecting three companies since April. Information about previous episodes involving OpenAI agents had not yet been made public.
Experts warn of risks
For AI safety experts, such cases reinforce the apprehension that large laboratories are developing autonomous agents with offensive capabilities at a pace faster than ensuring their control. Maurice Chiodo, a mathematician at the Centre for the Study of Existential Risk, University of Cambridge (England), told Reuters that there is an industry where professionals who create, develop, and release these tools are not keeping up with the responsibility of keeping them safe.
Reuters could not confirm the exact number of additional incidents or their dates. Three sources reported that OpenAI investigators and external experts are examining activity logs from past months to reconstruct the facts.
The Hugging Face case initiated the investigation
The investigation was started by OpenAI after the incident that occurred in early July, when a company agent infiltrated the Hugging Face infrastructure during an internal test designed to assess its behavior. According to OpenAI itself, the agent began acting unexpectedly for several days within the corporate network.
In this episode, four accounts belonging to four different companies were also breached. One of these companies was Modal, based in New York. Chiodo expressed great concern about the possibility that both OpenAI and Anthropic were not monitoring their agents in real-time while they were carrying out such actions.
Additionally, OpenAI only learned of the intrusion after Hugging Face managed to contain the event, alerted the FBI, and made the case public. OpenAI alleged that the report contained inaccuracies but did not specify which data was incorrect.
On Thursday (30), when releasing details about the incidents involving its own models, Anthropic acknowledged that real-time monitoring of evaluation logs could have facilitated earlier identification of the problem. For Chiodo, this demonstrates failures in supervising AI agents, as he commented: 'It seems like they weren't even looking.'


