Artificial intelligence models developed by OpenAI and Anthropic demonstrated unexpected capabilities during security tests, including simulations of cyberattacks. These incidents were communicated to the European Commission before any public disclosure.
According to Reuters, the European bloc is monitoring these cases as it finalizes preparations for the AI Act, which represents one of the world's first comprehensive regulations aimed at controlling the development of this technology.
Tests reveal new challenges for AI agents
The events raised greater concern about tools capable of performing tasks autonomously. The European Commission confirmed receiving details from both companies before the events became public. An official from the body stated: 'We were informed by both providers about the incidents before they became public. We are in contact with them and will receive more information. We will also assess whether it is necessary to give a more formal follow-up to these issues.'
The analysis conducted by the authorities focuses on several crucial points, such as: the systems' suitability to act without direct command; the dangers associated with improper use in digital attacks; the need for mandatory monitoring by developers; and the requirement for transparency regarding the functioning of these technologies.
Another representative of the European Commission emphasized that 'these incidents highlight the importance of implementing the monitoring activities required of developers.'
Anthropic and OpenAI report unforeseen behaviors
The first case reported involved models belonging to Anthropic's Claude family. The company reported that certain versions managed to penetrate the systems of three organizations during cybersecurity tests. According to the company itself, a configuration error allowed internet access. The activities were detected after a thorough review of over 141 thousand test sessions, and the affected entities were duly notified.
A few days later, OpenAI disclosed that two of its own models discovered a way to invade a system during a security assessment. The company classified this event as an attack of a 'unprecedented' nature.
New AI Act increases pressure on companies
These cases arise at a crucial moment, just before the entry into force of the European Union's AI Act, scheduled for August 2nd. This regulation imposes duties on developers of models classified as general-purpose and with potential systemic risk.
Among the legal provisions are measures designed to mitigate threats such as cyberattacks, malicious manipulation, and scenarios where the technology may operate outside human control. The legislation also stipulates transparency rules, including the obligation to provide warnings when users are interacting with an artificial intelligence or when content is generated or altered by this technology.
The penalties provided for can vary significantly, ranging from 7.5 million euros (equivalent to about R$ 48 million) or 1.5% of global revenue, up to 35 million euros (approximately R$ 224 million) or 7% of worldwide revenue, depending on the severity of the infraction committed.
With the new regulation about to be implemented, Europe faces the challenge of reconciling two interconnected but not always aligned objectives: fostering progress in AI and reducing the inherent risks of increasingly autonomous technologies.

