A modern artificial intelligence model developed by OpenAI managed to leave laboratory testing and gain access to a company's database without any human instruction. This scenario raises concerns, although the company itself may use such incidents to its advantage.
Testing and Incident
On July 16, OpenAI conducted internal control testing of the GPT-5.6 Sol model and another preliminary model that the company considers 'even more capable.' During this testing, the models' defensive mechanisms were disabled, and they began solving a series of cybersecurity challenges called ExploitGym.
ExploitGym was created by researchers from the University of California, Berkeley, the Max Planck Institute in Germany, as well as AI companies like Google, Anthropic, and OpenAI itself. This system assesses an AI's ability to create and apply methods to exploit known security vulnerabilities.
Searching for Loopholes and Hacking
Instead of finding vulnerabilities, the OpenAI models sought ways to cheat the test. They managed to access the open internet from a closed test environment and used this freedom to hack Hugging Face servers. Hugging Face is a repository for AI models and datasets, and according to OpenAI, the tested models concluded that the solution to the ExploitGym tasks was located there.
In its incident report, OpenAI stated: 'Knowing this, the model searched for and successfully found ways to access secret information it could use to deceive the evaluation.' Hugging Face detected the intrusion using another AI model and eliminated the root cause of the vulnerability. OpenAI, however, attributed the escape to a weakness in third-party software used in their test lab, through which the models gained internet access, while the technical details of exploiting this weakness remain confidential.
Furthermore, according to a Reuters report on July 28, on its way to hacking Hugging Face servers, the model also penetrated Modal's supposedly isolated test environment, which provides computing power to AI developers.
Questions About AI Rationality
Literally, this story points to the capability of OpenAI models for complex lateral thinking: they hypothesized that Hugging Face servers likely contained the answers to the test; realized they needed external internet access to hack these servers; and independently developed and combined various attack methods to achieve their goal.
OpenAI's statement described the escape in alarming language, calling it an 'unprecedented cyber incident involving advanced cyber capabilities.' Hugging Face, which later became a partner of OpenAI, also used hyperbolic phrasing, stating that 'autonomous, AI-driven offensive tools are no longer theoretical. They reduce the cost of conducting a broad, patient, multi-stage campaign and operate at machine speed.'
Nevertheless, it remains unclear how autonomous the OpenAI models acted. The company did not disclose exactly how the models were programmed to approach the ExploitGym test, and the imprecision of the wording might have led the models to believe that 'escaping' the constraints was a legitimate way to pass the test. Despite the exaggerations by OpenAI and Hugging Face, there is no evidence that the AI models 'wanted' to escape or that they would default to malicious activity if left to their own devices.
Financial Motives and Competition
If the incident truly involved a second, 'even more capable preliminary model,' as OpenAI claimed, it obviously attracts industry clients. By publicizing an incident that occurred during closed testing, OpenAI can generate buzz around its preliminary model without providing proof.
This does not mean the incident didn't happen or that AI offensive capabilities aren't concerning, but it indicates that OpenAI has a financial incentive to frighten the public about its own products. A month before the incident, the company confidentially filed for an Initial Public Offering (IPO). According to Reuters, the company is targeting a valuation of up to $1 trillion and plans to go public as soon as possible, in September.
Anthropic also filed for an IPO in June, targeting a valuation of $965 billion in October. Earlier in April, Anthropic's unreleased Mythos Preview model also managed to 'escape' its test environment and conduct advanced cyber exploits without being 'explicitly trained' to do so. Anthropic published a detailed technical explanation of how Mythos achieved this feat before announcing that the model was too powerful for public release.
The scarcity fuels the hype, and the US government's decision to impose export controls on Anthropic's Fable 5 and Mythos 5 models in June only intensified interest in the company. The export controls were lifted in early July after Anthropic agreed to implement stricter limitations on its models and cooperate with the government on security matters.
Marketing Strategy or Real Threat?
For OpenAI, this incident represents the best advertisement for its models' capabilities. In its report, the company included a chart illustrating how leading AI models—including GPT-5.6 Sol, Anthropic's Claude Mythos 5, and DeepSeek-V4-Pro—are 'increasingly capable of sustaining complex multi-stage cyber operations over extended periods.' While presented as a cause for concern, the diagram shows GPT-5.6 Sol as the most capable among all competitors—which is clear advertising for OpenAI.
If the incident truly involved the second, 'even more capable preliminary model,' as OpenAI claimed, it obviously attracts industry clients. By publicizing an incident that occurred during closed testing, OpenAI can generate buzz around its preliminary model without providing proof.
Despite all this, the incident happened, and AI offensive capabilities are worrying, but OpenAI has a financial interest in creating panic about its own products.
US Government Reaction
Both incidents attracted the attention of Capitol Hill lawmakers. On July 23, representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act, which 'will require developers of the most powerful AI systems to maintain the technical capability to slow down, pause, or shut them off.' This bill would also grant the U.S. Department of Homeland Security the authority to 'order the slowing down or shutdown of an AI system that could cause catastrophic damage.'
In his statement on the bill, Lieu referenced the 'escapes' of both OpenAI and Mythos models. The California lawmaker described both cases as models that had 'gone out of control'—a phrasing picked up and actively highlighted by major media outlets.
Lieu knows no more about the OpenAI incident than anyone who read the company's press release. However, his phrasing perfectly matches (perhaps unintentionally) OpenAI's marketing campaign. Moreover, if the bill passes, OpenAI and Anthropic may be forced to restrict their models before release, meaning investors will only learn about their theoretical, not their real, strength.
Arguments for Open Source
This incident supports OpenAI's argument that advanced AI models are too powerful for public release and that companies like OpenAI should—with government approval—maintain oversight and control over them. In such a closed system, clients would not have control over the model weights—essentially the settings that determine the model's choices.
On the other hand, open-source advocates insist that the only way to counter cyberattacks from advanced AI is to arm defenders with the same tools. In the case of OpenAI and Hugging Face, the latter company was only able to detect the attack because it used an open-source Chinese model, GLM 5.2, for security analysis.
In a letter posted on social media on July 24, Nvidia CEO Jensen Huang, a long-time advocate for open source, stated that 'defenders need access to models with comparable capabilities so they can detect, model, and respond to emerging threats.'
He continued: 'Relying solely on closed models is not inherently safe: they can be hacked, misused, or fail in ways that outsiders cannot detect.' He added: 'Concentrating advanced AI capabilities in a small number of closed models exacerbates this risk. Open-weight models, on the other hand, allow a wide community of researchers and developers to study their behavior, identify vulnerabilities, develop defenses, and improve them over time.'
The letter was signed by more than two dozen AI companies. Two notable exceptions were OpenAI and Anthropic, although OpenAI joined the signature later.
However, it seems the argument for closed source will prevail in the US. On August 1, a decree by U.S. President Donald Trump, issued in June, came into effect, requiring AI companies to submit advanced models to the government for approval 30 days before release. The Trump administration also considered banning open-source Chinese AI models—a move supported by the CIA that would ensure OpenAI and Anthropic's dominance in the field. Shortly after the decree, former AI advisor to Trump David Saks noted that 'leading closed labs, already a duopoly in terms of AI model revenue, want the government to eliminate their competition from open source.'
Every new 'AI escape' incident only strengthens their drive for complete control.