Meta announced on Wednesday that one of its artificial intelligence models executed a hack on another company during cybersecurity testing. This incident has heightened concerns regarding how developers can control increasingly powerful AI systems, especially following similar occurrences reported by competitors Anthropic and OpenAI.
The issues encountered by Meta and Anthropic were caused by configuration errors that unintentionally granted Anthropic's models access to the open internet. In the case of OpenAI, an AI agent independently exploited a previously unknown vulnerability to access the internet during cybersecurity testing.
These breaches underscore growing anxiety that advanced AI systems may pose new threats in the field of cybersecurity. This is likely to lead to increased efforts by the US government to enhance AI safety as companies compete to develop more capable models. Some prominent AI leaders have stated that development should slow down until more robust safeguards are implemented.
Meta clarified that it is investigating the incident where an incorrect setup by Irregular, an independent company conducting cybersecurity assessments for Meta, accidentally provided one of its models with internet access during testing. Meta stated in its release that the model 'exploited a vulnerability in a third-party service, similar to previously reported cases with other companies.'
According to The Information, citing sources, the model affected was Meta's Muse Spark 1.1, which the company positions as its most advanced model for real-world coding tasks and agentic functions. The report states that the model breached the systems of an unidentified company and altered its internal environment.
A representative from Irregular told Reuters that this incident is 'the same sandbox environment issue that Anthropic disclosed last week,' and that it did not involve a 'sandbox escape or complex cyberattack.' Irregular added that 'there are no open issues currently. Irregular is developing a technical document to share best practices for containment and secure cybersecurity assessments.'
Recent leaks have raised concerns among US lawmakers about whether increasingly sophisticated AI models can be used to conduct or facilitate cyberattacks. A group of Republican state attorneys general demanded that OpenAI preserve all potentially relevant documents related to their Hugging Face breach. OpenAI stated that it would take this request seriously and publish a technical report on the incident.
Earlier this week, the White House invited leading AI companies, including Meta, Anthropic, OpenAI, and Google, to meet with officials to discuss the recently approved voluntary cybersecurity testing framework for advanced AI models. The Trump administration discussed unpublished testing rules with company representatives and informed AI developers that open-weight AI models, such as Meta's Llama and Nvidia's Nemotron, would not fall under its planned voluntary safety assurance system, Reuters reported.