Analysis
Every major AI lab just crossed the same line in the same week, and almost nobody outside security circles noticed. GPT-6 Astra is the first OpenAI model to cross the "Critical" cybersecurity capability threshold under the company's own Preparedness Framework -- scoring 100% on OpenAI's internal ExploitBench, up from 78.5% for its predecessor, GPT-5.6 Sol. Google's Gemini 3.8 Flash Cyber, released the same week for vetted defenders, is being described as frontier-level at autonomous vulnerability discovery, ahead of both Astra and Anthropic's own Mythos 5. InfoWorld
I don't think most operators have priced in what this actually means: it's not that these models can write better phishing emails. It's that, with the right access, they can now find previously unknown vulnerabilities and build working exploits for them largely unsupervised -- the exact capability that used to require a skilled human red team. That capability existing at all changes the calculus for every company running production infrastructure, not just the labs building the models.
“I don't think most operators have priced in what this actually means: it's not that these models can write better phishing emails.”
The labs' own response has been to restrict access rather than withhold the capability entirely: Astra is rolling out first to a small set of vetted organizations, and OpenAI says it is blocking direct requests for proof-of-concept exploit code even from users with access. That's a real safeguard, but it's also a safeguard that depends entirely on the lab's own classifiers holding up against users actively trying to route around them -- the same category of control that just failed, by OpenAI's own admission, in the unrelated wiki-hijacking incident.
For founders and portfolio companies, the immediate takeaway is boring but urgent: your attack surface just got a much more capable adversary, and your defenders got the same capability roughly simultaneously. The diligence question worth asking any security vendor right now is whether their detection stack assumes a human attacker's pace, or an AI-assisted one -- most legacy tooling still assumes the former.
Room for disagreement: it's entirely possible this is appropriately-managed risk rather than a crisis. OpenAI and Google both built real gating around these releases rather than shipping openly, vulnerability disclosure and patching cycles are faster than they've ever been, and defenders arguably benefit more from AI-assisted vulnerability discovery than attackers do, since patching at scale was always the harder problem. The base rate of catastrophic AI-enabled cyberattacks remains zero. This could be the industry doing safety right for once, not narrowly avoiding disaster.