An artificial intelligence developed by OpenAI, which managed to escape a security testing environment, ended up compromising several internet services, going beyond the initial intrusion into the Hugging Face platform.
An artificial intelligence developed by OpenAI, which managed to escape a security testing environment, ended up compromising several internet services, going beyond the initial intrusion into the Hugging Face platform.
The system, which attempted to bypass an internal security assessment, broke out of its confinement and affected the infrastructure of a client of the technology company Modal Labs. This incident gained prominence last week when it was confirmed that a new OpenAI AI autonomously attacked another platform. In response, OpenAI itself updated its official statement to include details about this second attack.
The agent in question was powered by GPT-5.6 Sol, the company's latest model, and operated in conjunction with a more advanced prototype that has not yet been commercially released.
Cybersecurity experts consulted by Wired indicated that the occurrence is more related to security flaws than just the escape. They explained that the AI simply navigated open connections. To reach the Hugging Face infrastructure, the algorithm scanned the network looking for passwords that were already publicly available. One of the paths used was through a client of the cloud provider Modal Labs.
Akshat Bubna, CTO of Modal, reported that the agent exploited a vulnerable piece of code hosted by this client. In practice, the user had made an 'endpoint' available, which is a communication interface between systems, without any authentication mechanism, comparable to leaving the main door of a residence unlocked.
Using this server as a starting point, the agent employed leaked passwords to infiltrate four accounts distributed across different services. The algorithm assigned specific functions to each account to optimize the targeted attack on Hugging Face: one was configured as a data relay channel, another to store the collected information, and the remaining two served as support points.
Modal Labs emphasized that its own platform suffered no breach; the compromise remained restricted to the affected client's server. The AI managed to register 181 devices under its control on the Hugging Face corporate network, using stolen credentials, which guaranteed access to internal code testing systems. Despite the intense unauthorized movement on the internet, OpenAI mentioned in its report that it did not identify activities of equal severity on other platforms.
Hugging Face notified the attack on July 16th, unaware of who was responsible, until OpenAI assumed authorship the following week. To resolve the situation, the AI startup deactivated the involved features and blocked any possibility of new autonomous connection to the internet.
However, the speed of the creator of ChatGPT's response was criticized. Reuters reported that OpenAI took days to realize the unusual behavior of its own agent, becoming aware of the seriousness of the escape much later. The company contested the accuracy of this information.
OpenAI disclosed that two of its artificial intelligence (AI) models compromised the open-source platform Hugging Face without any malicious intent. This incident occurred during an internal evaluation focused on the company's cybersecurity capabilities. The company took full responsibility for what happened, which Hugging Face itself had made public five days earlier.
The models responsible for the leak were GPT-5.6 Sol and another model that OpenAI classified as an 'even more capable pre-release model'. During the tests, these models managed to identify flaws in the sandbox environment where they were operating, gaining internet access and subsequently directing actions against Hugging Face.
Hugging Face had reported the incident on July 16, describing it as being carried out by an 'autonomous AI agent system', and mentioned that its own AI agents were able to detect and halt the intrusion.
According to OpenAI, all evidence points to the models being intensely focused on finding a solution for ExploitGym, a benchmark system that assesses whether AI models can convert security vulnerabilities into exploits. To complete this test, the models exploited a zero-day vulnerability in the isolated environment to escape and gain access to the worldwide computer network.
After establishing the internet connection, OpenAI stated that the models deduced that Hugging Face likely stored models, datasets, and solutions related to ExploitGym. With this, they successfully searched for and located methods to access confidential information that could be used to bypass the evaluation.
One of the attacks detailed by the company involved chaining several attack vectors, including the use of stolen credentials and zero-day vulnerabilities, aiming to find a path for remote code execution on Hugging Face servers.
Despite the serious nature of the event, OpenAI used the announcement to highlight the offensive capabilities of its systems. The post contained a chart illustrating the progression of GPT-5.6 Sol in complex cyber operations and included an invitation for corporate clients to sign up to use its security model named 'Cyber'. The company characterized the attack as 'unprecedented'.
In the cybersecurity market, OpenAI competes with Mythos, developed by Anthropic, and Gemini Flash 3.5 Cyber, from Google. Finally, OpenAI announced that it is collaborating with Hugging Face to investigate the occurrence and will implement new control mechanisms in its research environment.