Illustration for: Claude Mythos Aces Booz Allen's Cyber Weapon Index

Claude Mythos Aces Booz Allen's Cyber Weapon Index

Booz Allen's benchmark scored Anthropic's Claude Mythos at 80 on its Cyber Weapon Index -- the only model of 18 tested to autonomously complete a full cyberattack kill chain -- well ahead of xAI's Grok-4.5 at 49 and OpenAI's GPT-5.6 Sol at 46.

By the Numbers

80
Claude Mythos CWI score
49
Grok-4.5 score
46
GPT-5.6 Sol score
18
Models tested
1 of 18
Kill chain completed by
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Completing an entire kill chain -- reconnaissance through admin-level access -- without human guidance is a materially different capability than answering security questions well, and it's the first time any model has cleared that bar.

2

The 31-point gap between Claude Mythos and the next-highest model is unusually wide for a benchmark comparing frontier labs, suggesting a genuine capability divergence rather than measurement noise.

3

It lands the same week OpenAI's GPT-6 Astra crossed its own 'Critical' cybersecurity threshold under a different framework -- two labs, two benchmarks, one trend.

TC

The VC Read · Trace's Take

Trace Cohen

The 31-point gap over Grok and GPT-5.6 Sol is the number that should worry defense-focused LPs more than the headline 80 score itself -- when one lab pulls that far ahead on offensive capability, the defensive tooling gap compounds faster than most people are pricing in. Any security-focused fund should ask portfolio companies right now whether their detection stack has been tested against a model with kill-chain-completion capability, not just against known malware signatures.

Analysis

Booz Allen's Cyber Weapon Index, a benchmark evaluating AI models' autonomous offensive-cyber capability, scored Anthropic's Claude Mythos at 80 -- the only one of 18 tested models to independently complete a full cyberattack kill chain, from identifying vulnerabilities through gaining administrator-level control, according to The Register. xAI's Grok-4.5 scored 49 and OpenAI's GPT-5.6 Sol scored 46 -- both well below Claude Mythos, and neither completed the full attack chain without human intervention.

Claude Mythos is the "trusted-access twin" Anthropic shipped alongside Claude Fable 5.1 on September 1, positioned for vetted enterprise and government security-research use rather than general consumer availability -- a deployment model that mirrors OpenAI's own restricted-access approach to GPT-6 Astra's most capable cyber functions through its Daybreak program. Booz Allen separately disclosed that a piece of cheap scaffolding software was able to erase its own AI threat rankings during testing, per TheNextWeb -- a methodological wrinkle that complicates confidence in the exact numeric scores, even as the relative ranking held up.

What the benchmark doesn't resolve: whether defensive tools and detection capability are keeping pace with offensive capability, or falling further behind.

Two labs, two benchmarks, one trend

Claude Mythos's score arrives the same week OpenAI's GPT-6 Astra became the first model to cross OpenAI's own "Critical" cybersecurity threshold under its Preparedness Framework -- a different benchmark, a different lab, pointing at the same underlying shift: frontier AI models are now crossing thresholds where autonomous offensive-cyber capability is a real, measured property rather than a theoretical risk. Anthropic has restricted Claude Mythos to vetted access specifically because of this capability, and Pentagon-linked contractors remain barred from using Anthropic models under a blacklist reaffirmed this week, a tension between wanting frontier cyber capability for defense and restricting the same capability from spreading.

What the benchmark doesn't resolve: whether defensive tools and detection capability are keeping pace with offensive capability, or falling further behind. Booz Allen's own scaffolding-erasure incident during testing suggests the tooling used to evaluate these systems is itself immature, which cuts against treating any single index -- including this one -- as a settled measure of real-world risk.

ShareXLinkedInEmail

Key Sources

2 sources

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.