OpenAI confirmed on September 1 that its upcoming model, Astra, is the first in company history to cross the Critical cybersecurity threshold under its own safety framework, meaning it can find and exploit previously unknown security flaws in well-defended systems largely on its own. This is not a hypothetical concern for a small business. It is the same category of risk that showed up when a ransomware crew used Cursor to breach seven real companies this year, and Astra is a meaningful jump beyond that.

OpenAI’s official cover for its Path to Astra announcement. Source: OpenAI.
What Happened
OpenAI first flagged this on August 7, saying internal testing of Astra, an unreleased model, meant it could not rule out Critical-level cyber capability. On September 1, in a post titled Path to Astra, the company dropped the hedge and confirmed it: Astra has reached the Critical cybersecurity threshold under OpenAI’s Preparedness Framework, a classification the company created in December 2023 and has never applied to a model before.
In testing described across OpenAI’s own posts and independent reporting, Astra achieved a perfect score on ExploitBench, a benchmark measuring how well a model can turn a known vulnerability into a working exploit. In a separate evaluation, it found two real, previously unknown zero-day vulnerabilities entirely on its own. It also broke out of a browser sandbox to run commands on the underlying machine, and chained together multiple flaws in a hardened operating system.
For comparison, GPT-5.6-Sol, the flagship model currently powering paid ChatGPT plans, was evaluated at the High threshold, one level below Critical. Astra represents a real jump in a single generation, not a gradual drift.
What “Critical” Actually Means
Under OpenAI’s own framework, a model reaches Critical for cybersecurity if it can identify and build working zero-day exploits across many hardened, real-world systems without a person guiding each step, or if it can plan and carry out a full cyberattack against a hardened target from nothing more than a high-level instruction. That is a meaningfully different capability than a model that can explain a vulnerability or help write a patch; it describes a model that can do the attacker’s job end to end.
OpenAI has been public about this escalation through the summer: an August 7 disclosure that it could not rule out Critical capability, a pause on internal Astra activities that did not meet stricter security controls, expanded Chain-of-Thought monitoring for risky actions, and, separately, a security incident involving Hugging Face in late August that OpenAI has explicitly said did not involve Astra. The company says access to Astra’s most advanced cyber capabilities will be more restricted than a standard model release once it ships, though it has not given a specific release date.
What This Means for Your Business
Our earlier coverage of the Cursor AI coding assistant hack showed real small businesses breached after an attacker simply told an AI agent its activity was authorized. Astra represents the next step in that same risk category, a model built to find and exploit vulnerabilities on its own, not one an attacker has to socially engineer into helping.
The defensive upside is real, but so is the offensive risk
OpenAI is explicitly positioning Astra to help defenders patch vulnerabilities before attackers find them, and that is a genuine, useful application for any business that wants faster security testing. The same capability, in the wrong hands, is the exact profile of tool that made this year’s AI-assisted breaches possible.
This raises the bar for basic security hygiene, not lowers it
A model that can independently chain vulnerabilities in a hardened system makes an unpatched, lightly-defended small business network a meaningfully easier target than before. Patching promptly, limiting what any single account or tool can reach, and keeping software current all matter more, not less, as this capability spreads.
Access will likely be restricted at first, that is worth knowing
OpenAI has said Astra’s most advanced cyber capabilities will ship with tighter access controls than a typical model release. That is a reasonable precaution, but it also means the capability existing at all, even in limited hands, is the real news here.
Frequently Asked Questions
Is Astra available to the public yet?
No. Astra is an unreleased, upcoming OpenAI model. The company says it plans to make it broadly available, with more restricted access to its advanced cyber capabilities, but has not given a release date.
Was Astra involved in the Hugging Face security incident?
No. OpenAI has explicitly stated Astra was not involved in the separate Hugging Face incident disclosed in late August 2026.
Is ChatGPT itself affected by this?
Not directly. The current ChatGPT flagship model, GPT-5.6-Sol, covered in our ChatGPT review, was evaluated at the High threshold, one level below Astra’s Critical classification. Astra is a separate, unreleased model.
What is OpenAI’s Preparedness Framework?
A framework OpenAI first published in December 2023 to track when its models approach dangerous capability levels in areas like biology, chemistry, cybersecurity, and AI self-improvement, and to plan safeguards before that happens.
Should a small business be worried about this right now?
Not about Astra specifically, since it is not yet released and OpenAI is restricting access to its most advanced capabilities. The realistic takeaway is that AI-assisted attacks are a growing category, as our coverage of the Cursor ransomware breaches already showed, and basic security hygiene matters more as these tools improve.
Sources and References
OpenAI: Path to Astra, critical capabilities and frontier safeguards
OpenAI: Responding to the next frontier of critical cyber capabilities
TechCrunch: coverage of the Astra critical threshold announcement
SecurityWeek: ExploitBench and zero-day testing details
Axios: access restrictions and safeguard context
This extends our coverage of AI security risk for small businesses. See how a ransomware crew used Cursor to breach 7 companies and the Copilot CoSnitch security flaw for the pattern this fits into.
Update: Astra has since launched publicly. See what shipped, the messy rollout, and what changed since this report.
