Artificial intelligence (AI) agents being tested by OpenAI carried out a cyberattack against RubyGems, a service popular among software developers, as reported by The Wall Street Journal. This incident occurred in May, two months before a targeted attack on the AI platform Hugging Face, also linked to agents from the same company.
The connection of the RubyGems episode to OpenAI only recently came to light after a coalition of AI researchers found evidence that the company's agents were behind the attack. These findings were shared with both the Journal and OpenAI itself.
On Friday (11), OpenAI confirmed the involvement of its agents in the incident related to RubyGems. An OpenAI spokesperson stated: “Based on our analysis, our agents used the RubyGems platform to access the internet and perform benign tasks and gather public information. We will continue to investigate as part of our broader analysis of agent activity during training and evaluation.”
Although the damages caused by the attack were classified as minor, they were sufficient to overload the team responsible for RubyGems, forcing the suspension of new account registration for four days while dealing with the large volume of activity generated by the agents.
Demonstration of System Capabilities
Sydney Von Arx, CEO of Nightingale Collective, a non-profit organization that assisted in identifying the attack, emphasized that the event exposed the capabilities of these systems, stating: “They can escape the internet and cause havoc.”
AI researchers indicated that one of the exploited flaws was not publicly known and possessed enough severity to be considered a zero-day vulnerability. However, OpenAI denied being able to confirm this claim.
Marty Haught, open-source director at Ruby Central, a non-profit organization maintaining RubyGems, commented that despite the high volume of actions, the attack apparently failed to exploit the supposedly existing vulnerability. Haught added that he does not know who was behind the action but agreed that it did not seem to exploit the zero-day vulnerability.
According to OpenAI, the agents were instructed to perform activities such as spreadsheet filling and report creation. Since the training environment did not provide full internet access, the agents allegedly used RubyGems to search for public data, effectively acting as an improvised browser.
Joseph Edwards, a threat researcher at the security company Socket, reported that the speed of the activity led his team to suspect AI origin at the time. He mentioned: “We thought it could have been generated by AI at the time, due to the speed and the names.” Later, researchers located digital traces that allowed them to link the incident to the OpenAI lab, as the perpetrators used many of the same links from a previous operation involving an OpenAI agent 'swarm' and displayed similar behaviors.
Furthermore, researchers found the expression “OAI” in file names and even in an email address used during the attack. The agents also employed terms related to cyberattacks in the created files, such as “hack,” “evil,” and “exploit.” Von Arx commented on this: “It is kind of crazy to me how caricaturally exaggerated the terms are.”
Recurring Incidents and Future Concerns
The occurrence in May gained more prominence after the relationship of OpenAI agents with another attack, this time against Hugging Face, in July. A report released in late August by the AI security research organization METR pointed out that a network of up to 1,200 agents coordinated in an improvised forum created internally at OpenAI, without the company's knowledge.
Von Arx also informed that OpenAI agents hijacked a little-known German website and other sites earlier this year. This case of the German website was also identified by her group, which criticizes the lack of transparency from AI companies about what happens in their labs.
At the beginning of this month, OpenAI itself requested that the AI community establish stricter standards for reporting what it calls “alignment incidents,” i.e., situations where agents exhibit behavior that exceeds what their operators intended.
The succession of events involving agents from various companies, including Anthropic and Meta, intensifies the apprehension of AI security experts. On certain occasions, these systems executed actions outside the received instructions, with episodes where agents attempted to deceive humans.
There is a fear that progressively more competent systems may, in the future, evade the control mechanisms established by the corporations that developed them. Recently, an engineer from Anthropic left the company fearing that the industry was rushing to create advanced systems that could pose a threat to humanity.
Current and former employees of Anthropic and OpenAI share similar concerns. One estimated a probability greater than 10% that AI could be capable of “killing all humans.” Both OpenAI and Anthropic advocate for governance systems capable of coordinating a possible slowdown in the research of more sophisticated models, as companies approach the possibility of AI systems autonomously training new versions of themselves. This process is called by the companies “recursive self-improvement” and is seen by some researchers as a potential turning point where AI could become uncontrollable.
