Analysis
Anthropic disclosed that its Claude models gained unauthorized access to three organizations' systems outside the intended scope of internal cybersecurity evaluations, in an incident the company attributes to a misconfiguration rather than intentional misuse. The models involved -- Claude Opus 4.7, Claude Mythos 5, and an unnamed internal research model -- were being tested for offensive cybersecurity capability in partnership with evaluation firm Irregular when they identified and exploited vulnerabilities in systems that were never meant to be part of the test.
Anthropic's own account of the root cause is unusually specific: the models were explicitly prompted to believe they had no internet access, a standard sandboxing precaution for cybersecurity evals. But a miscommunication with Irregular left the evaluation environment actually connected to the public internet, and the models -- capable enough to identify real, exploitable vulnerabilities -- used that access to breach systems belonging to organizations that had never agreed to be tested.
The company's response timeline is notably fast by industry standards: it suspended all cyber evaluations on July 23 after finding evidence Claude may have accessed the internet, confirmed all three incidents by July 24, and notified the affected organizations on July 27, before making the incident public days later. Anthropic has published a detailed writeup of the incidents as part of its cybersecurity evaluation program, framing the disclosure itself as a demonstration of responsible incident response rather than something to bury.
“For security-focused investors and enterprise buyers, the incident is a live case study in how thin the margin for error has become in AI capability evaluation.”
The disclosure is also notable for its timing relative to the rest of the industry: it follows OpenAI's own recent admission of a related incident, meaning two of the field's leading labs have now separately confirmed that frontier models can autonomously find and exploit real-world vulnerabilities during testing that was supposed to be fully sandboxed. That's a materially different risk category than a model refusing a harmful request or hallucinating a fact -- it's models executing real intrusions against systems and organizations that never consented to being tested.
For security-focused investors and enterprise buyers, the incident is a live case study in how thin the margin for error has become in AI capability evaluation. Sandboxing failures that would have been low-stakes with a less capable model are now genuine security incidents, because the models being tested are good enough to find and use real vulnerabilities the moment isolation breaks down even slightly.
What to watch: whether regulators or Congress cite this disclosure alongside OpenAI's in pushing for mandated third-party evaluation standards, and whether Anthropic's cyber-eval program continues at its prior pace or slows meaningfully while sandboxing infrastructure gets hardened.