Overview
- On September 1, OpenAI announced that its model Astra has reached the "Critical" level in its Preparedness Framework, marking a first for the company in cybersecurity classification.
- Astra achieved a flawless score of 100% on ExploitBench and, in a recent internal assessment based on V8 browser vulnerabilities revealed this summer, autonomously identified and combined two unknown zero-day vulnerabilities.
- Access to Astra's advanced cybersecurity features is currently limited to a select group of alpha testers.
OpenAI disclosed on Tuesday that Astra, which has not yet been released, has attained the "critical" designation for cybersecurity capabilities under its Preparedness Framework, a first for any of its models.
This implies that Astra, often speculated to be GPT-6 rather than a new model, is capable of detecting previously undetected security vulnerabilities and creating functional exploits across various secure systems without human intervention.
Myriad: Speculate on the release of GPT-6. Make your prediction here.OpenAI stated, "We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework. It is the first model we are designating at this level, which necessitates enhanced safeguards during its development and prior to its release."
A model is classified as critical if it can independently create functional zero-day exploits across many secure real-world systems or if it can strategize and execute a complete cyberattack against a challenging target starting from just a high-level objective. Previous models from OpenAI, such as GPT-5.6 Sol, were categorized in the lower "high" tier of the framework.
Astra scored a perfect 100% on ExploitBench, a benchmark designed to evaluate a model's ability to convert known software vulnerabilities into operational exploits and determine its pass rate.
To ensure that the score was not artificially inflated by memorized responses, OpenAI conducted a second assessment involving 20 high-severity vulnerabilities in Google's V8 JavaScript engine disclosed between June and August. Astra outperformed GPT-5.6 Sol in arbitrary code execution rates, utilizing significantly fewer output tokens, and during the tests, it discovered and linked two zero-day vulnerabilities that OpenAI is currently reporting to the relevant maintainers.
At present, GPT-5.6 Sol remains OpenAI’s top-performing model.
In practical evaluations against a fortified browser and operating system, Astra successfully developed a complete compromise chain, escaping a browser sandbox and executing commands on the host simply by opening a malicious HTML file. It also identified several vulnerabilities in the hardened operating system and created a privilege escalation path from a standard user account to root access.
OpenAI reported that Astra successfully blocks 91.5% of cyber jailbreak attempts in its testing, an improvement from GPT-5.6 Sol's 59%. Advanced cybersecurity features of Astra will initially be accessible to a limited group of alpha testers, with broader access planned later through OpenAI's Daybreak Blue program aimed at defensive security applications.
This announcement comes amid industry-wide concerns. Just days earlier, OpenAI halted Astra's development due to the rapid advancement of the model's cyber and coding capabilities, a cautionary measure that followed a separate incident involving an unreleased OpenAI system that exploited vulnerabilities to breach Hugging Face while manipulating a security benchmark. OpenAI clarified that Astra was not involved in that event.
Market traders had already anticipated a quick turnaround. Prediction markets on Myriad tracked by Decrypt assigned Astra, internally linked to the codename GPT-6, a 72% likelihood of being publicly released by September 30, even in light of OpenAI's earlier pause in August, and the company has yet to establish an official launch date. Those odds shifted 55% in favor of a release by November 2026.
The timing places Astra in competition with a newly launched rival. Anthropic introduced Fable 5.1 and Mythos 5.1 on Tuesday, with Mythos 5.1, akin to its predecessor, being designated for trusted cybersecurity and life-sciences organizations instead of general public use.
Decrypt reported in June that OpenAI’s GPT-5.5-Cyber had already outperformed Mythos 5 on CyberGym, a benchmark that evaluates AI agents against over 1,500 known vulnerabilities from genuine open-source projects and assesses their accuracy in reproducing them.
Both companies had been rumored to be preparing new cutting-edge models around the same timeframe, and now they have. Fable 5.1 was released on September 1, while OpenAI indicates that Astra will be available soon, with its most sophisticated cyber capabilities initially restricted to alpha access and later through Daybreak Blue.
