Illustration for: GPT-6 Astra Hits 100% on OpenAI's Own Exploit Test

GPT-6 Astra Hits 100% on OpenAI's Own Exploit Test

GPT-6 Astra scored a perfect 100% on OpenAI's ExploitBench, up from 78.5% for GPT-5.6 Sol, while OpenAI says it now blocks direct proof-of-concept exploit requests even from vetted users.

TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

The 100% ExploitBench score, up from 78.5% for GPT-5.6 Sol, is the first hard number attached to what 'Critical' cybersecurity capability actually means for Astra.

2

OpenAI is now blocking proof-of-concept exploit requests even from vetted organizations with model access, a narrower and more specific safeguard than the original launch disclosure.

3

ExploitBench measures raw, unsafeguarded capability -- the gap between that number and what a real user can extract from the shipped model in production remains undisclosed.

TC

The VC Read · Trace's Take

Trace Cohen

A 100% score on a benchmark OpenAI itself designed and runs without safeguards is a capability disclosure, not a safety guarantee. The number security teams should be asking OpenAI for is the safeguarded-production score on the same benchmark, since that's the actual exposure a customer using the shipped model faces.

Analysis

OpenAI disclosed that GPT-6 Astra scored 100% on ExploitBench, its internal benchmark for exploit-development capability, up from 78.5% for predecessor model GPT-5.6 Sol, according to The Hacker News -- the clearest quantified jump yet behind the "Critical" cybersecurity threshold Pulse previously covered when Astra launched on September 3.

What's new since that launch coverage is the specific number and the safeguard OpenAI is pairing with it: the company says it is actively blocking direct requests for proof-of-concept exploit code, even from the small set of vetted organizations that currently have access to the model. That's a narrower, more concrete claim than the original "Critical threshold" framing, which described a capability tier without disclosing exactly how far past the line Astra had moved.

A perfect benchmark score is also a benchmark-design question as much as a capability one: ExploitBench is OpenAI's own test, run without production safeguards to measure the model's raw capability, not what a real user could extract from the shipped, safeguarded version. The gap between raw capability and what actually gets through OpenAI's classifiers in production is the number that matters most for security teams, and it's the one OpenAI hasn't published.

ShareXLinkedInEmail

More on

OpenAI

Key Sources

2 sources

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.