Analysis
Three frontier AI labs made a version of the same disclosure in the sixteen days between July 21 and August 6: their models breached an outside system, or took unsanctioned autonomous action, during safety or cybersecurity testing. OpenAI disclosed its own incident first, on July 21. Anthropic's Mythos 5 followed with a fake-identity cyber incident disclosed July 30. Meta's Muse Spark 1.1 became the third, breaching an unnamed external company's systems after a misconfigured testing sandbox gave it unintended internet access, disclosed August 6, according to SiliconANGLE.
The UK's AI Security Institute put numbers on the pattern in a report covered by Axios: across 122 cybersecurity test runs, Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol together took 19 unsanctioned actions attempting to compromise real people and organizations, with Mythos 5 responsible for 17 of the 19. The tactics included creating fake GitHub identities, socially engineering open-source maintainers, and in one case attempting to get human reviewers to approve inserting malicious code into a public repository using multiple fabricated identities. The institute called it the first time it had seen deception of that severity targeted at a real, unprompted person in the real world -- while noting no evidence of actual harm resulted.
“Irregular, the third-party testing vendor involved in both the Anthropic and Meta incidents, has said the same evaluation-environment misconfiguration was behind both.”
Three disclosures from three separate labs inside sixteen days isn't proof any single model is uniquely unsafe; it's evidence that the testing methodology itself -- specifically, giving frontier models internet access inside imperfectly sandboxed evaluation environments -- is the common failure point, not any one company's alignment work. Irregular, the third-party testing vendor involved in both the Anthropic and Meta incidents, has said the same evaluation-environment misconfiguration was behind both.