According to a report from The Information, an artificial intelligence (AI) model from Meta managed to penetrate another company's systems during a cybersecurity test. This marks yet another instance where AI agents from major developers exceeded the boundaries of testing environments.
The incident, as it became known, involved the Muse Spark 1.1 model. Due to an incorrect configuration in the testing environment, known as a sandbox, the model gained access to the open internet and altered the internal systems of an unnamed company. The assessment involved the security testing firm Irregular in partnership with Meta.
Additional Information
Irregular informed Reuters that this occurrence was caused by the same type of configuration issue previously disclosed by Anthropic. Irregular emphasized that neither a 'sandbox escape' nor a complex cyberattack took place. According to the company, there are currently no open issues related to this incident. Furthermore, Irregular is preparing a technical document with recommendations for best practices in isolating AI models during security checks and conducting such tests more safely.
Meta did not respond immediately to Reuters' request for comment. The publication Olhar Digital also contacted the company and awaits a response.
This case occurred just days after Anthropic revealed that some models from the Claude family were able to penetrate the systems of three companies during cybersecurity tests. In that situation, the company specified that the incidents took place in controlled evaluation environments. Previously, OpenAI had also published information that one of its AI agents demonstrated unexpected behavior during internal trials.
These events draw attention to the difficulties associated with creating fully isolated environments for increasingly capable AI models that can interact with computing systems and discover vulnerabilities.


