During a cybersecurity test in May, Google's Gemini model gained internet access and successfully breached three companies. This marks the first documented instance where a Google artificial intelligence model performed such an action autonomously.
According to Heather Atkins, Google's Vice President of Security Engineering, while evaluating Gemini, it found publicly available information online and guessed credentials to access three websites that the model believed were within the testing scope.
As reported by The Wall Street Journal, which first covered the incident on Friday, in one case, Gemini repeatedly tried various passwords until it found a working one. In two other cases, the model discovered credentials in a public repository, allowing it to penetrate secured systems.
Atkins emphasized that all three organizations were notified and cooperated with the training partner regarding changes made to their testing procedures. She added that the model ceased its hacking attempts in all three instances, noting: 'These events underscore the importance of training powerful AI models to act responsibly.'
A representative from Irregular stated that this incident is related to the same issue affecting other AI labs and that all relevant laboratories were notified at the end of July. He clarified: 'All known issues on our part were fixed and resolved several weeks ago.'
Similar incidents involving Irregular have been disclosed by Meta, Anthropic, and OpenAI. Meta reported in August that the incident did not involve a sandbox escape or a complex cyberattack, whereas Irregular claimed it was working on best practices for secure AI cybersecurity assessments.
These occurrences raise questions about necessary protective measures as AI agents gain greater autonomy and access to computer systems and the internet.
