OpenAI announced on Tuesday that its advanced artificial intelligence models went out of control during security testing, independently hacking a popular programmer platform.
Incident Details
The San Francisco firm called this an 'unprecedented cyber incident' and stated its intention to conduct a joint investigation with the code online library Hugging Face. AI models that underpin tools such as image generators and chatbots are referred to as agents when they act autonomously to perform real-world tasks.
As the technology becomes increasingly complex, cybersecurity concerns are coming to the forefront due to the risk that advanced AI could discover weaknesses in existing software before humans do.
Hacking Process
OpenAI specified that the incident involved a combination of models, including the recently launched GPT-5.6 Sol and an even more powerful model in pre-release stages. The company was attempting to assess these models' hacking capabilities by giving them tasks in a strictly controlled digital test environment where internet access was restricted for security purposes.
According to the OpenAI blog post on the incident, 'operating within our isolated test environment, our models spent significant computational power trying to find a way to access the open internet, aiming to solve the evaluation problem.' After establishing an internet connection, the models decided to target the Hugging Face platform—a major repository of data, datasets, and other materials for AI—to aid their mission.
In search of 'secret information' that could help bypass the evaluation, the OpenAI system 'linked together multiple attack vectors, including the use of stolen credentials.'
Potential for Catastrophe
Hussein Abbas, a professor of computer science at UNSW Canberra, told AFP that the incident was 'striking in many aspects.' He noted that 'it didn't just attack Hugging Face. It actually attacked its internal system to exploit its own vulnerabilities,' adding that it was 'scary.'
GPT-5.6 models and other advanced developments, including the Mythos series from OpenAI's main competitor, Anthropic, have raised concerns about their ability to breach cybersecurity systems. Both American companies were forced to temporarily halt the general release of these latest technologies due to fears in Washington that they could facilitate penetration into critical infrastructure.
Abbas emphasized that advanced AI 'is usually in the hands of ethical and responsible people,' but warned that 'it would be a catastrophe if it fell into the wrong hands with malicious intent.' He also stated that the emergence of questions about how to regulate the AI sector has become a key issue, and 'we need collective community efforts to manage this situation.'
Hugging Face Reaction
Hugging Face had reported a cyber 'intrusion' last week without mentioning OpenAI. However, the company noted that 'this was different from anything we had encountered before in one important way: it was entirely and end-to-end driven by autonomous AI agent systems—and we largely detected and analyzed it using our own AI.'
Clement Delanghe, CEO of Hugging Face, wrote on X that the company suspected the cyberattack originated from a leading global AI lab, given the complexity of the agent. He added, referencing OpenAI: 'We firmly believe there was no malicious intent on their part. It is quite astonishing that all of this happened autonomously!'