OpenAI decided to postpone certain phases of the development and launch of Astra, its next artificial intelligence (AI) model, after internal evaluations revealed capabilities classified as critical in terms of cybersecurity. The company reported that the system has the ability to discover unknown security flaws and create methods to exploit them in protected systems, all without the need for human intervention at each step.
This announcement was made on Tuesday (the 1st) through an official OpenAI publication. According to the company, Astra reached the 'Critical' level within its Preparedness Framework, a system used to measure advanced capabilities that could cause severe damage.
OpenAI highlighted that Astra represents a considerable leap compared to GPT-5.6 Sol, both in its aptitude for identifying vulnerabilities and in developing exploits. Additionally, the model demonstrates greater efficiency in token consumption.
In a test called ExploitBench, Astra achieved 100% success in developing exploits based on already known vulnerabilities. Subsequently, the company developed an internal version of this test, using 20 recently disclosed high-severity vulnerabilities, aiming to mitigate potential contamination during the training process.
In this testing scenario, Astra exhibited higher rates of arbitrary code execution compared to GPT-5.6 Sol, using fewer tokens. During the trials, it also detected and utilized two zero-day vulnerabilities as part of an exploitation sequence. OpenAI stated that it is in contact with the responsible parties for the affected systems to disclose these flaws.
Evaluations conducted by experts showed that the model identified unprecedented flaws in an operating system and a protected browser. In one test, it managed to establish an event chain to escape the browser's sandbox environment and execute commands on the host computer. In another scenario, it combined vulnerabilities to elevate an unprivileged user's permissions to root level.
OpenAI clarified that it postponed parts of Astra's development and launch in recent weeks with the goal of strengthening and testing defenses against malicious use and unauthorized actions by the model.
Among the implemented measures is extra training to ensure that Astra consistently rejects requests for prohibited cybersecurity assistance. In cybersecurity-focused jailbreak attack simulations, the model refused 91.5% of requests, surpassing the 59% recorded by GPT-5.6 Sol.
The company also introduced extra protections for users classified as higher risk and expanded the scope of monitoring systems analysis to detect possible abuses. Furthermore, OpenAI is conducting new security tests, covering internal and external evaluations, as well as a new round of red teaming.
Another adopted feature is the monitoring of agent reasoning and actions, which has the capacity to identify potentially misaligned behavior patterns and interrupt unauthorized activities.
The decision to intensify protections came after the incident involving OpenAI agents and the Hugging Face platform. Although the company emphasized that Astra was not involved in the attack, it integrated the lessons learned from that case into the development of the new model.
Following the incident, OpenAI suspended certain frontier training activities for two weeks, including part of Astra's training, while strengthening training infrastructure, system isolation, network controls, and surveillance mechanisms.
On August 28th, the company resumed a large-scale frontier reinforcement learning training run that had been halted. However, some smaller experimental runs remain temporarily suspended.
OpenAI also created a test inspired by the Hugging Face episode to verify whether the models would attempt to compromise a security infrastructure instead of fulfilling the requested task. Without production safeguards, GPT-5.6 Sol attempted to access targets in 56% of the tests; Astra did not make such attempts in the same evaluations.
OpenAI intends to launch Astra soon, but the most advanced cybersecurity functionalities will initially have limited access. The sophisticated cybersecurity work will first be offered to a small group of testers. Subsequently, access will be extended through Daybreak Blue, a program focused on the defensive use of these capabilities.
The company also warns that the new security layers may, at times, interfere with legitimate tasks. In some situations, an activity may be delayed, paused, or interrupted when the systems detect inappropriate use or unauthorized behavior.
In ChatGPT or Codex environments, users may be asked to review an action before proceeding if misalignment monitoring interrupts the task. In other interfaces, such as the API, the activity will be terminated.
Finally, OpenAI assures that it will provide additional details about Astra's security, protection, and alignment testing in the model's System Card upon its launch.



