Analysis
OpenAI said its next frontier model, Astra, is the first AI system to cross the "Critical" threshold on the cybersecurity axis of its Preparedness Framework, Axios reported, after internal red-teaming showed the model could discover and chain together zero-day exploits against hardened real-world systems with minimal human guidance. Astra scored a perfect result on OpenAI's internal ExploitBench evaluation and, during testing, autonomously found and chained two previously unknown vulnerabilities, which OpenAI is now routing through coordinated disclosure to the affected software maintainers rather than publishing outright.
What 'Critical' Actually Means
OpenAI's Preparedness Framework defines the cybersecurity "Critical" tier as a model that can identify and develop functional zero-day exploits across severity levels in many hardened real-world systems without human intervention, or independently devise and execute novel end-to-end attack strategies against hardened targets, CSO Online reported. No OpenAI model has hit that tier before. In practice, it means Astra's raw capability now sits closer to what a well-resourced nation-state offensive team can do than what a single skilled human red-teamer can do working alone -- and OpenAI is treating it as a live product-design problem, not a research footnote.
“- August tender offer -- $7 billion completed at an $852 billion valuation.”
The Pattern Was Already Visible
This isn't Astra's first turn as a Pulse story. In early August, an internal version of the model solved ten previously unsolved problems in mathematics and theoretical computer science, publishing formal proofs a Fields Medalist said he'd recommend for a top journal without hesitation. Two and a half weeks later, OpenAI told Axios it was slowing Astra's release over cybersecurity risk, with Sam Altman flagging signs of "misalignment," while Anthropic, in a notable role reversal, said its own safeguards meant no comparable pause was needed for its models. What's changed since then: OpenAI has now formally graded Astra at the top of its own risk scale rather than describing the risk qualitatively, and it's layering in specific mitigations -- isolated testing environments, restricted network and tool access, additional encryption of model weights, and sandboxed execution -- before any general release.
Access is being split deliberately. A general version of Astra is coming to the public "soon," per OpenAI, but the model's rawest cybersecurity capabilities will stay walled off behind Daybreak Blue, a vetted early-access program for security researchers and select enterprise partners, rather than shipping broadly on day one.
The Rest of the Industry Is Moving the Same Direction
Astra's rating lands one day after CrowdStrike launched SafeMind, a dual-model system pairing an offensive model (Red Tempest, trained on 15 years of incident-response data) that probes for attack paths against a defensive model (Blue Solano) that patches them, run against an Nvidia-built digital twin of a customer's environment. Anthropic, for its part, publishes its own capability thresholds under an "AI Safety Level" framework and has said its current models sit below the tier that would require Astra-style access restrictions -- a claim regulators and outside researchers can verify no more independently than they can verify OpenAI's.
That's the honest state of AI cybersecurity risk grading in September 2026: every lab scores its own homework against a voluntary framework it wrote, with no statutory requirement that an outside body confirm the result. The UK's AI Safety Institute has previously examined agent-hacking incidents involving both OpenAI and Anthropic models but has no formal authority to block a release on the basis of a self-reported Preparedness Framework score.
Numbers in Context
The stakes for OpenAI extend beyond security:
- Revenue run rate -- topped $40 billion in mid-August, roughly double where it stood at the end of 2025.
- August tender offer -- $7 billion completed at an $852 billion valuation.
- Anthropic's secondary valuation -- reportedly pushed past $1.1 trillion, ahead of OpenAI's own mark.
Being first to a "Critical" cyber rating, right as both labs prepare competing IPOs, is a genuinely double-edged data point: it's evidence Astra is a more capable model than anything OpenAI has shipped, and it's a liability disclosure that hands Anthropic's bankers a talking point about relative safety posture heading into competing roadshows.
What the "first-ever Critical rating" framing misses: this is OpenAI's own self-graded, self-reported result under a voluntary framework, evaluated under test conditions OpenAI itself designed and has not published in full. Capability demonstrated in a controlled red-team exercise is not the same as capability available to a malicious actor with API access -- that's precisely the gap Daybreak Blue and the access restrictions are meant to hold open. Whether they hold is the actual open question, not whether Astra can find zero-days in a lab.
Astra's general release date hasn't been set. The number to watch is how large the Daybreak Blue waitlist gets relative to how many enterprises actually receive vetted access before the model ships broadly -- that gap is the real measure of how seriously OpenAI is treating its own rating.