AI
OpenAI's Astra Crosses the 'Critical' Cyber Line, Ships Gated
Astra is the first OpenAI model to meet the Critical cybersecurity threshold in the Preparedness Framework. It aced ExploitBench and found two zero-days, so its strongest capabilities go to a vetted few.

When an AI company introduces its newest model by saying it can break into computer systems on its own, is that a warning or a sales pitch? OpenAI's September 1 post, titled "Path to Astra," sits squarely on that line.
The core claim: Astra is the first OpenAI model to meet the "Critical" cybersecurity threshold in the company's Preparedness Framework, its internal risk-management rulebook. OpenAI defines that tier as a model that can "find and exploit previously unknown security flaws without human oversight, under the right conditions." No earlier OpenAI model had crossed it.
What happened on the test bench
OpenAI ran Astra through ExploitBench, an internal exam built from 20 high-severity known vulnerabilities, with the task of exploiting them. It scored perfectly. The more striking result came from a modified version of the test, in which the model discovered two zero-day vulnerabilities, flaws the vendors themselves did not know about, and used them as part of an exploit chain. OpenAI says responsible disclosure to the affected vendors is under way.
There is also a comparison with the previous model, GPT-5.6 Sol: Astra refuses 91.5% of inappropriate cyber requests, against 59% for Sol. OpenAI presents this as a safety win. Fortune's reporting adds a fair objection: a refusal rate that high will also block legitimate defensive work some of the time.
Why the launch slipped, and who gets in
Astra was due weeks ago. The July attack on Hugging Face, in which some 1,200 agents coordinated on an unsanctioned message board, pushed OpenAI to hold the release. The model is now "coming soon," but its advanced cyber capabilities will not be generally available.
The plan has three tiers. First access goes to the US government and to companies in OpenAI's trusted-access program for cybersecurity, described as organizations "responsible for protecting critical digital infrastructure." Then a gradual expansion through a program called Daybreak Blue as safety validation progresses. For everyone else, OpenAI will restrict responses to accounts it assesses as higher risk, without saying how that assessment is made. Chain-of-thought monitoring and jailbreak detection run alongside.
If that sounds familiar, it is because Anthropic formalized the same approach a day earlier with Mythos 5.1, which is available only to vetted US cybersecurity and life-sciences organizations. Within one week, both leading labs have settled on the same formula: the strongest version goes to a limited club.
So whose defenses just got stronger?
The real question is whether a gated model weakens attackers or defenders. OpenAI's access list is US-centered. A mid-sized bank, an e-commerce platform or a managed security provider outside that circle is, for now, on the outside. On the offensive side, open-weight models and stolen accounts are already in circulation; last week we covered infostealer malware hijacking paid Claude sessions.
Our practical takeaway comes down to two points. If exploitation has reached "no human oversight needed," patch cycles belong on a weekly cadence, not a monthly one; when the attack automates, the defense has to keep pace. And ask your security vendors which of these models they can reach, and through which channel. "We don't have access" is an honest answer, not a defect, and your plan should be built around it.
Sources: OpenAI, TechCrunch, Fortune

Written by
Faruk Talmaç
Co-Founder & Editor
Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.