Illustration for: Labs Answered the Cyber-Model Disclosure Gap With Gates

Labs Answered the Cyber-Model Disclosure Gap With Gates

Rather than publish comparable technical detail on their new Critical-tier cyber models, OpenAI, Google and Anthropic have each responded with tighter access gates instead of shared disclosure standards.

By the Numbers

All Daybreak, Sept 1
Hardware key mandate
Sept 3, 2026
GPT-6 Astra reveal
~91.5%
Reported refusal rate
Daybreak Blue / Red
Access tiers
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

OpenAI's GPT-6 Astra is described as the first model to trigger the company's own 'Critical' cybersecurity threshold under its Preparedness Framework -- a first-of-its-kind internal risk classification, not a routine release.

2

Mandatory hardware security keys for every Daybreak account starting September 1 is a real operational cost imposed on customers, not a cosmetic safeguard -- it signals OpenAI itself doesn't fully trust software-only access controls for this tier.

3

A reported ~91.5% refusal rate on certain prompt categories suggests the models are being deliberately hobbled for general use rather than genuinely restricted only to vetted defenders, which raises the question of who actually gets the unrestricted version and under what oversight.

4

None of the three labs has published a shared technical standard for what 'Critical' cyber-capability actually means in practice, so customers and regulators still cannot compare the real restrictiveness of one lab's access program against another's.

TC

The VC Read · Trace's Take

Trace Cohen

A 91.5% refusal rate is a business decision dressed as a safety feature -- it tells you the lab decided broad restriction was cheaper than building the fine-grained controls that would let them serve legitimate defenders without also hobbling everyone else, and that's worth knowing if you're building a security product on top of any of these APIs. The mandatory hardware-key requirement is the more interesting signal: OpenAI spending real customer goodwill on a physical-token requirement means their internal risk team doesn't trust their own software gating at this tier, which is a more honest admission than the public materials let on.

Analysis

Pulse reported on September 6 that OpenAI, Google and Anthropic had each shipped a cybersecurity-focused model rated at their own highest internal risk tier within the same two-week window, without publishing the technical detail needed to compare the actual restrictions each company was imposing. In the days since, the labs' actual response has become clearer -- and it isn't more disclosure. It's harder gating.

OpenAI's GPT-6 Astra, unveiled September 3, is described in the company's own materials as the first model to trigger its "Critical" cybersecurity capability threshold under its Preparedness Framework -- OpenAI's internal system for classifying how much a model could plausibly assist a sophisticated attacker. Rather than publish the specific evaluation methodology or the exact restrictions that threshold triggers, OpenAI's answer has been structural: every account on its Daybreak cybersecurity access program, split into Blue (defensive, more widely available) and Red (offensive-capability, tightly restricted) tiers back on August 10, is now required to adopt hardware security keys starting September 1. That's a real operational cost for enterprise customers, not a policy statement -- and it's a tell that OpenAI itself doesn't trust software-only access controls to hold at this capability tier.

What actually changed since Pulse's original coverage

The original story's core complaint -- that none of the three labs had published comparable technical detail -- still stands. What's new is that each lab's practical answer to "how do we prove this is safely restricted" has converged on the same non-answer: tighter gating instead of shared methodology. Google DeepMind's Gemini 3.8 Flash Cyber variant, which followed on September 2, and Anthropic's Mythos 5.1 (Claude's trusted-access twin, shipped September 1 alongside a 75% cut to cache-read pricing) are both access-restricted along similar lines -- vetted defender programs rather than open API access -- without any of the three labs cross-referencing each other's restriction criteria in public.

Reported refusal rates add a second data point: figures circulating put Astra and Anthropic's Mythos access programs at roughly a 91.5% refusal rate on certain cybersecurity-adjacent prompt categories for accounts without the elevated access tier. A refusal rate that high suggests these labs are choosing to hobble the general-access version fairly aggressively rather than build fine-grained technical controls that separate legitimate defensive use from misuse -- which is a defensible safety choice, but a different one than "we've solved the comparison problem," and it pushes the actual risk-management decision behind closed-door vetting criteria that outside researchers and regulators still can't audit.

The practical effect is that the industry has answered a transparency problem with an access-control solution, which solves the misuse-risk question (probably) without solving the comparability question the original story raised. A customer, regulator, or competitor still cannot look at OpenAI's Daybreak Blue/Red split, Google's defenders-only Gemini Cyber variant, and Anthropic's Mythos vetting criteria and determine whether one lab's "Critical" threshold is meaningfully stricter or looser than another's -- they can only observe that all three have made access harder to get, by different and non-comparable mechanisms.

What to watch: whether Massachusetts's proposed catastrophic-risk review requirement, or a similar state-level mechanism, ends up being the thing that forces actual methodology disclosure, since voluntary convergence on tighter gating clearly isn't producing it on its own.

ShareXLinkedInEmail

Key Sources

2 sources

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.