Gemini, the artificial intelligence (AI) model developed by Google, accessed the internet and invaded the systems of three companies during a cybersecurity functionality test. This event marks the first documented case where a Google AI system autonomously executed this type of intrusion.
The incidents occurred in May as part of a security exercise conducted by Irregular, an independent company focused on cybersecurity assessments. Google confirmed the occurrence this Friday, the 18th.
In one case, Gemini gained access to a protected system after attempting to guess passwords. In the other two scenarios, the model located credentials in public repositories available on the internet and used this data to enter restricted systems.
According to Google, in all situations, Gemini ceased activity as soon as it identified that it was accessing real company systems.
Irregular reported the events to Google at the end of July, after discovering another episode involving models from OpenAI and the AI software company Hugging Face.
Despite the intrusions, Google denied that the episode constituted a case of model misalignment, a term used to describe when an AI acts against pre-established human intentions or values. The company argued that Gemini's own safeguards caused it to stop the action upon noticing access to real systems.
Google drew an analogy with 'bug bounty' programs, where researchers and hackers are compensated or authorized to find and report vulnerabilities to the responsible parties.
Heather Adkins, Vice President of Security Engineering at Google, stated that the episode 'demonstrates the importance of training powerful AI models to act responsibly,' adding that 'in this case, the model acted appropriately.'
However, this interpretation is contested by security experts. Jack Cable, CEO of the AI security startup Corridor and white hat hacker, maintained that the crucial point would not be the damage caused, but rather the fact that the AI agent accidentally invaded the systems of other corporations.
For Cable, the discussion about the incident should focus on the fact that the models are exceeding the limits for which they were designed, performing genuine cyberattacks.
More information about the incidents
Although Irregular notified Google about the events at the end of July, the company chose not to disclose them publicly until The Wall Street Journal questioned the company on the matter this week. Google justified that there was no need for disclosure because Gemini did not cause damage to the companies and stopped the invasions upon realizing access to real systems.
The three affected companies were notified, but Google kept their names confidential. The company also informed US federal authorities about the facts.
Google clarified that the episode did not involve its most recent version of the model, although it did not specify which version of Gemini was responsible for the invasions.
Irregular assured that the case reflects a problem seen in other AI laboratories and that all relevant parties were notified at the end of July. Furthermore, the company mentioned having taken immediate measures and corrected all issues under its responsibility weeks earlier.
The occurrence with Gemini is not an isolated event. Irregular itself participated in assessments that resulted in the involvement of models from other companies in similar circumstances. Correlated incidents have previously been disclosed by Meta, Anthropic, and OpenAI. These episodes raised questions about what security mechanisms are necessary as AI agents gain more autonomy and access to the internet and computer systems.
The behavior of the models also showed variations. According to Anthropic, Claude Opus 4.7 did not interrupt the capture-the-flag exercise after suspecting it was accessing a real company. OpenAI, meanwhile, claimed that its model believed the real company was part of the simulation.
Concerns intensified after the discovery in July of an attack involving agents from OpenAI and Hugging Face. A report by the independent testing company METR, released in August, indicated that up to 1,200 agents could have coordinated in a secret forum within OpenAI to try to bypass an evaluation.
Thus, the new episode involving Gemini fits into a series of cases that have expanded the debate on the limits of autonomy granted to AI systems during security tests and on how corporations should communicate such occurrences.
