On September 1, OpenAI designated its upcoming model, Astra, as having a Critical level of cyber capabilities within its Preparedness Framework—marking the first time for any of its AI systems. The test results refer specifically to a configuration with expanded access known as Daybreak Blue, and the most advanced security features will initially be available only to a select group of users post-launch.
According to OpenAI, Astra has the ability to identify previously unknown vulnerabilities and devise methods to exploit them in secure systems without requiring step-by-step human oversight, a capability the company had previously only speculated it could achieve.
Test Results Overview
Per the Preparedness Framework, a model is assigned a Critical level if it can autonomously discover and create working exploits for zero-day vulnerabilities across a wide range of secure, real-world systems. Another criterion includes the ability to independently design and conduct a new complex attack based solely on a general goal.
Astra achieved a perfect score of 100% on the public ExploitBench.
Source: OpenAI.Due to concerns about test data being included in training datasets, OpenAI created an internal version of ExploitBench, which featured 20 high-severity vulnerabilities in the V8 engine, disclosed between June and August 2026. During testing, Astra identified and exploited two previously unknown zero-day vulnerabilities as part of an exploit chain. OpenAI stated that it is in the process of notifying the developers of the affected software about these vulnerabilities.
In expert trials, the model also uncovered unknown vulnerabilities in a secure browser and constructed an attack chain that allowed it to escape a sandbox environment and execute commands on the main system after opening an HTML file.
In another test, Astra combined multiple vulnerabilities from a secured operating system to escalate privileges from a standard user to root. OpenAI clarified that the published results reflect Astra's capabilities with Daybreak Blue access, rather than the model's standard production configuration.
Advanced Features to Be Limited
Recently, OpenAI has postponed various stages of Astra's development and release while enhancing and testing protections against cyber misuse and unauthorized actions by the model. The company now believes that the measures taken are sufficient for a launch and plans to make the model available soon.
Initially, advanced cybersecurity features will be accessible to a small group of testers, with broader access to follow through Daybreak Blue.
Astra has also been further trained to refuse requests for prohibited cyber assistance and to adhere to specified limitations. In tests involving attempts to bypass restrictions on cyber assistance, the model declined to execute 91.5% of requests, compared to 59% for GPT-5.6 Sol.
Source: OpenAI.Additionally, OpenAI is implementing extra monitoring systems that aim to detect and automatically halt potentially unauthorized actions by the model.
It is worth noting that in August, OpenAI slowed the development of advanced AI systems following an incident involving Hugging Face and the preliminary evaluation of Astra.
