Analysis
OpenAI paused every internal Astra activity that doesn't meet a strengthened set of security controls after evaluations showed the unreleased model's autonomous coding and cybersecurity capabilities had grown large enough that the company can no longer rule out a "Critical" risk rating -- the highest tier in its Preparedness Framework and the first time OpenAI has applied that designation to any model, according to CNBC and Bloomberg. A "Critical" rating specifically means the model could plausibly launch cyberattacks against sophisticated defenses autonomously, without a human prompt specifying how.
Astra now operates under isolated testing environments with restricted network and tool access, encrypted model weights, sandboxed execution, and chain-of-thought monitoring designed to interrupt high-risk activity in real time, according to Forbes. Government agencies and selected third-party safety organizations get evaluation access before any public release. OpenAI says it still intends a broad release once Astra satisfies its safety and security requirements, but has not given a timeline.
“Government agencies and selected third-party safety organizations get evaluation access before any public release.”
The pause lands in a specific context: Pulse has tracked a summer of AI-related security incidents, including reports that OpenAI and Anthropic models breached other companies' cybersecurity defenses during internal testing between July 9 and 13. That backdrop is part of why House Democrats are now pushing for OpenAI and Anthropic executives to testify before Congress. OpenAI's decision to publicly disclose the Critical-tier concern, rather than quietly delaying Astra without explanation, reads as an attempt to get ahead of that regulatory pressure by demonstrating its own safety process is working as designed.
The self-reporting is also the thing worth being skeptical of. OpenAI is both the entity running the safety evaluation and the entity that decides what counts as sufficiently mitigated to lift the pause -- there's no independent, binding third-party check on when Astra's risk tier gets downgraded back to something releasable. The Preparedness Framework is OpenAI's own internal policy, not a government-mandated standard, which means the pause is voluntary in the same sense the White House's own AI safety framework is voluntary: it holds only as long as OpenAI decides it should.