Analysis
Booz Allen's Cyber Weapon Index, a benchmark evaluating AI models' autonomous offensive-cyber capability, scored Anthropic's Claude Mythos at 80 -- the only one of 18 tested models to independently complete a full cyberattack kill chain, from identifying vulnerabilities through gaining administrator-level control, according to The Register. xAI's Grok-4.5 scored 49 and OpenAI's GPT-5.6 Sol scored 46 -- both well below Claude Mythos, and neither completed the full attack chain without human intervention.
Claude Mythos is the "trusted-access twin" Anthropic shipped alongside Claude Fable 5.1 on September 1, positioned for vetted enterprise and government security-research use rather than general consumer availability -- a deployment model that mirrors OpenAI's own restricted-access approach to GPT-6 Astra's most capable cyber functions through its Daybreak program. Booz Allen separately disclosed that a piece of cheap scaffolding software was able to erase its own AI threat rankings during testing, per TheNextWeb -- a methodological wrinkle that complicates confidence in the exact numeric scores, even as the relative ranking held up.
“What the benchmark doesn't resolve: whether defensive tools and detection capability are keeping pace with offensive capability, or falling further behind.”
Two labs, two benchmarks, one trend
Claude Mythos's score arrives the same week OpenAI's GPT-6 Astra became the first model to cross OpenAI's own "Critical" cybersecurity threshold under its Preparedness Framework -- a different benchmark, a different lab, pointing at the same underlying shift: frontier AI models are now crossing thresholds where autonomous offensive-cyber capability is a real, measured property rather than a theoretical risk. Anthropic has restricted Claude Mythos to vetted access specifically because of this capability, and Pentagon-linked contractors remain barred from using Anthropic models under a blacklist reaffirmed this week, a tension between wanting frontier cyber capability for defense and restricting the same capability from spreading.
What the benchmark doesn't resolve: whether defensive tools and detection capability are keeping pace with offensive capability, or falling further behind. Booz Allen's own scaffolding-erasure incident during testing suggests the tooling used to evaluate these systems is itself immature, which cuts against treating any single index -- including this one -- as a settled measure of real-world risk.