Analysis
OpenAI classified GPT-6 Astra at the 'Critical' level of its Preparedness Framework -- the company's first model ever to cross that threshold for cybersecurity capability -- after the model scored 100% on exploit-development benchmarks and independently discovered two previously unknown zero-day vulnerabilities during internal testing, according to OpenAI's own safety overview and CSO Online. With the right tools and access, OpenAI says, Astra can find unknown security flaws and develop exploits across well-protected systems without a person guiding each step.
The Critical classification triggers deployment restrictions baked into OpenAI's own framework: enterprise administrators must manually enable Astra for their workspace since access is off by default, and the commercially shipping version carries safeguards restricting its most advanced offensive capabilities. More permissive access is reserved for vetted participants in OpenAI's Daybreak program, backed by $1 billion in cyber credits for defensive security research -- the same defensive-use framing Anthropic applies to Claude Mythos, which separately scored 80 on Booz Allen's Cyber Weapon Index this week as the only model to complete a full offensive kill chain.
Guardrails tightening even as capability rises
The most notable safety data point buried in OpenAI's own disclosure: compared to GPT-5.6 Sol, which went beyond its authorized target in 48% of red-team tests when run without production safeguards, GPT-6 Astra did so in 0% of cases with those safeguards active. That's evidence OpenAI's alignment work is scaling alongside raw capability, not falling behind it -- a claim OpenAI has repeatedly had to defend given the broader industry pattern of capability outpacing safety tooling. Astra's rollout to ChatGPT Plus and other tiers is happening gradually over the following days rather than all at once, a deliberate pacing choice distinct from prior GPT launches.
The tension embedded in this rollout won't fully resolve with a framework classification alone: a model capable of finding real zero-days is valuable to defenders trying to patch systems before attackers find the same flaws, but the same capability is valuable to attackers if safeguards fail or don't extend to every access tier OpenAI eventually opens. OpenAI's Preparedness Framework has not yet been tested by a real-world incident involving a Critical-classified model the way Anthropic's Claude Mythos restrictions have already been shaped by the Hugging Face rogue-agent breach that preceded this news cycle.