Analysis
OpenAI released GPT-6 Astra on Thursday, the first model the company has designated "Critical" under its own Preparedness Framework -- a rating that means Astra can autonomously discover previously unknown security flaws and build functional exploits against well-defended systems without a human directing every step, CNBC reported. OpenAI said it delayed parts of the model's development over the past several weeks specifically to strengthen and test protections against cyber misuse before proceeding with the rollout Pulse covered as a delay two days ago.
The launch came with OpenAI's boldest public framing yet of its own progress. President Greg Brockman, asked directly whether Astra represents artificial general intelligence, told reporters "I think it might be about this model," then closed the briefing with: "Welcome to the AGI era." That's a deliberately unresolved claim -- Brockman didn't say Astra definitively is AGI, he said it might be, which lets OpenAI stake out the framing in headlines without committing to a testable definition anyone could hold the company to later.
Gated Access, Familiar Playbook
Astra's rollout follows the same gated architecture Pulse has now tracked across three frontier labs this week:
- OpenAI Astra -- Daybreak program partners (vetted cybersecurity users) get access first; ChatGPT Plus, Pro, Business, Enterprise, the OpenAI API and Amazon Bedrock follow within days. Standard API pricing runs $10 per million input tokens and $50 per million output, with a Fast mode at roughly double that.
- Google Gemini 3.8 Flash Cyber -- restricted to the Fairwind Program for vetted defenders, Pulse covered today.
- Anthropic Mythos 5.1 -- restricted to vetted cybersecurity and life-sciences partners, while Fable 5.1 ships broadly.
Astra's published benchmarks are aggressive by design: it saturates FrontierMath Tier 4 at 97.6% and ExploitBench at 100%, and tops OSWorld 2.0 computer-use tasks at 72.6% while completing them roughly 47% faster than OpenAI's prior Sol model, The New Stack reported. The marquee ARC-AGI-3 score of 99.9% is the headline number driving the AGI framing, but it depends on a stateful, expensive evaluation harness -- ordinary stateless API calls, the way most developers will actually use the model, score far lower on the same benchmark.
For AI application developers, Astra's coding and computer-use benchmarks reset the bar competitors have to clear -- but the pricing, at $10/$50 per million tokens for standard mode, is also a meaningful step up from prior-generation frontier pricing, meaning the capability jump comes with a real cost increase for any startup building Astra into a production product at scale. VCs evaluating AI-application portfolio companies should model both the benchmark improvement and the unit-economics hit before assuming a model upgrade is a pure win.
The real risk the "AGI era" framing overstates is how much of the 99.9% headline number is testing conditions rather than product reality: a stateful harness that isn't how the model ships to paying customers inflates a benchmark score in a way that doesn't reflect what most users will experience. Brockman's own hedge -- "I think it might be about this model" -- is itself a limitation worth noting: it's an admission that OpenAI isn't claiming AGI has definitively arrived, just that it's comfortable letting the ambiguity do work the capability claims might not fully support on their own.
The more concrete story is the Critical cyber rating, not the AGI framing -- a model that can autonomously find and exploit unknown vulnerabilities is now shipping to a gated set of customers, and the real test of OpenAI's safeguards comes the first time Daybreak access gets misused or leaked rather than the first time someone calls the model AGI on a benchmark chart.