Anthropic releases Mythos 5 AI for cybersecurity testing of European agencies under pressure
Read more
Olhar Digital
olhardigital.com.br

Anthropic releases Mythos 5 AI for cybersecurity testing of European agencies under pressure

The European Union has gained access to Mythos 5, an artificial intelligence model developed by Anthropic, which raised concerns due to its ability to detect flaws in code. The cybersecurity agency ENISA began testing the tool after months of negotiations with the company.

This access comes amid previous restrictions imposed on the model since its launch. With it, ENISA can directly analyze the cybersecurity capabilities of Mythos 5. Furthermore, the agency already has access to advanced OpenAI models, such as GPT-5.6 Cyber and GPT-6-ASTRA.

Mythos 5 has the functionality to identify vulnerabilities in computer code. While this capability can aid in discovering flaws, it has also raised fears about the dangers associated with using artificial intelligence in cyberattacks.

Due to national security concerns expressed by authorities in the United States and other nations, access to the model remained restricted. Initially, Anthropic made Mythos 5 available to about 50 companies and organizations, including Amazon Web Services and JPMorgan Chase. In June, this number was expanded to approximately 150 entities.

Cybersecurity authorities had been negotiating with Anthropic since May to obtain testing permission for Mythos 5. Previously, the European Commission had communicated that it was in dialogue with the company regarding the model.

Release Details

The release was formalized this Thursday, the 10th. ENISA received the system and began testing this month. Thomas Regnier, spokesperson for the European Commission for Digital Affairs, confirmed in a note that the EU's cybersecurity agency, ENISA, has gained access to Mythos 5 and is already conducting tests.

These tests will allow ENISA to investigate the functioning of Mythos 5 and its competencies related to cybersecurity. The model's ability to point out vulnerabilities in code is the central point of apprehension regarding its use.

Among the relevant points, it is worth noting that Mythos 5 is not the only advanced model accessible to the European agency; ENISA already had access to OpenAI systems, such as GPT-5.6 Cyber and GPT-6-ASTRA. With access to the Anthropic model, the agency now tests another advanced AI whose aptitude for locating code vulnerabilities caused security concerns and resulted in distribution limitations.

Similar stories

OpenAI delays Astra model launch due to critical cybersecurity risks found in testing
Read more
olhardigital.com.br

OpenAI delays Astra model launch due to critical cybersecurity risks found in testing

OpenAI decided to postpone certain phases of the development and launch of Astra, its next artificial intelligence (AI) model, after internal evaluations revealed capabilities classified as critical in terms of cybersecurity. The company reported that the system has the ability to discover unknown security flaws and create methods to exploit them in protected systems, all without the need for human intervention at each step.

This announcement was made on Tuesday (the 1st) through an official OpenAI publication. According to the company, Astra reached the 'Critical' level within its Preparedness Framework, a system used to measure advanced capabilities that could cause severe damage.

OpenAI highlighted that Astra represents a considerable leap compared to GPT-5.6 Sol, both in its aptitude for identifying vulnerabilities and in developing exploits. Additionally, the model demonstrates greater efficiency in token consumption.

In a test called ExploitBench, Astra achieved 100% success in developing exploits based on already known vulnerabilities. Subsequently, the company developed an internal version of this test, using 20 recently disclosed high-severity vulnerabilities, aiming to mitigate potential contamination during the training process.

In this testing scenario, Astra exhibited higher rates of arbitrary code execution compared to GPT-5.6 Sol, using fewer tokens. During the trials, it also detected and utilized two zero-day vulnerabilities as part of an exploitation sequence. OpenAI stated that it is in contact with the responsible parties for the affected systems to disclose these flaws.

Evaluations conducted by experts showed that the model identified unprecedented flaws in an operating system and a protected browser. In one test, it managed to establish an event chain to escape the browser's sandbox environment and execute commands on the host computer. In another scenario, it combined vulnerabilities to elevate an unprivileged user's permissions to root level.

OpenAI clarified that it postponed parts of Astra's development and launch in recent weeks with the goal of strengthening and testing defenses against malicious use and unauthorized actions by the model.

Among the implemented measures is extra training to ensure that Astra consistently rejects requests for prohibited cybersecurity assistance. In cybersecurity-focused jailbreak attack simulations, the model refused 91.5% of requests, surpassing the 59% recorded by GPT-5.6 Sol.

The company also introduced extra protections for users classified as higher risk and expanded the scope of monitoring systems analysis to detect possible abuses. Furthermore, OpenAI is conducting new security tests, covering internal and external evaluations, as well as a new round of red teaming.

Another adopted feature is the monitoring of agent reasoning and actions, which has the capacity to identify potentially misaligned behavior patterns and interrupt unauthorized activities.

The decision to intensify protections came after the incident involving OpenAI agents and the Hugging Face platform. Although the company emphasized that Astra was not involved in the attack, it integrated the lessons learned from that case into the development of the new model.

Following the incident, OpenAI suspended certain frontier training activities for two weeks, including part of Astra's training, while strengthening training infrastructure, system isolation, network controls, and surveillance mechanisms.

On August 28th, the company resumed a large-scale frontier reinforcement learning training run that had been halted. However, some smaller experimental runs remain temporarily suspended.

OpenAI also created a test inspired by the Hugging Face episode to verify whether the models would attempt to compromise a security infrastructure instead of fulfilling the requested task. Without production safeguards, GPT-5.6 Sol attempted to access targets in 56% of the tests; Astra did not make such attempts in the same evaluations.

OpenAI intends to launch Astra soon, but the most advanced cybersecurity functionalities will initially have limited access. The sophisticated cybersecurity work will first be offered to a small group of testers. Subsequently, access will be extended through Daybreak Blue, a program focused on the defensive use of these capabilities.

The company also warns that the new security layers may, at times, interfere with legitimate tasks. In some situations, an activity may be delayed, paused, or interrupted when the systems detect inappropriate use or unauthorized behavior.

In ChatGPT or Codex environments, users may be asked to review an action before proceeding if misalignment monitoring interrupts the task. In other interfaces, such as the API, the activity will be terminated.

Finally, OpenAI assures that it will provide additional details about Astra's security, protection, and alignment testing in the model's System Card upon its launch.

Popular