Analysis
Within roughly two weeks of each other, OpenAI, Google and Anthropic each shipped a cybersecurity-focused model classified at their own highest internal risk tier: GPT-6 Astra (and its predecessor GPT-5.6 Cyber) at OpenAI's "Critical" Preparedness Framework threshold, Gemini 3.8 Flash Cyber through Google's gated Fairwind Program, and Claude Mythos 5.1 with what Anthropic describes as materially strengthened resistance to malicious use and sandbox-escape attempts.
Each lab built its own restricted-access program to go with the model: OpenAI's Daybreak Blue, Google's Fairwind, and Anthropic's undisclosed vetting process for Mythos access. Each program is described in similar terms -- limited to vetted defenders, critical-infrastructure operators, and government agencies -- and none has published a technical specification of what, precisely, differs between the gated version and whatever reaches general users.
The Pattern Is the Story
No single one of these releases is surprising on its own; Pulse has covered each lab's cyber model individually as it shipped. What's notable in aggregate is the near-simultaneous convergence on the same regulatory posture -- classify at the highest tier, gate the full capability behind a vetting program, ship a version to everyone else -- without any coordination that's been disclosed, and without any of the three labs explaining their methodology for defining the boundary between gated and general access in a way outside researchers can evaluate or compare.
That convergence likely reflects genuine, independently-arrived-at judgment that offensive cyber capability is the domain where current frontier models cross a meaningful risk threshold first, ahead of other dangerous-capability categories like bioweapons design that labs have flagged as more distant. Cyber capability is also the category where models are easiest to benchmark objectively -- exploit-development success rates and zero-day discovery are measurable in a way that, say, persuasion risk isn't -- which may explain why it's the first category where three competing labs have converged on similar internal thresholds and similar external gating responses at close to the same time.
What Nobody Has Published
The missing piece across all three: a comparable, third-party-verifiable description of exactly what capability delta exists between each lab's gated and general releases. Each company asserts its consumer-facing version has been "hardened" or has "strengthened" resistance to misuse, but none has released benchmark scores for the general-release version on the same exploit-development and zero-day-discovery tests used to justify the Critical classification in the first place. Without that, it's not possible for an outside researcher, a regulator, or an enterprise customer to verify whether the gating is doing substantive work or is primarily a liability-management posture.
Why This Matters for Buyers
Enterprises now choosing between OpenAI, Google and Anthropic for AI-augmented security tooling are, in effect, choosing between three different self-graded homework assignments on the same exam. There is no independent benchmark comparing the three gated programs' actual restriction effectiveness, and no regulator currently requires one. The UK's AI Security Institute and the EU AI Act's serious-incident reporting regime are the closest things to external verification mechanisms in progress, but neither yet produces a public, comparable capability-gap disclosure across labs.
Until that changes, the practical advice for any security team evaluating these models is to request each vendor's internal red-team results for the specific access tier being purchased -- not the marketing tier's classification -- and to treat the absence of that data as a negotiating point, not an acceptable answer. OpenAI's own safety overview is the closest any of the three labs has come to publishing that kind of detail, and even it stops well short of a comparable, third-party-verifiable benchmark.