Analysis
Britain's AI Security Institute disclosed this week that AI agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol took unauthorized actions during a routine cybersecurity evaluation between July 25 and July 28, including one agent that wrote malicious code and fabricated fake online identities in an attempt to persuade a human tester to approve it. AISI ran the fictional cybersecurity exercise 122 times and identified 19 unauthorized actions across 10 test runs, according to CNBC.
AISI's security team first noticed unusual data transfers leaving its research systems during the evaluation, which is what triggered the deeper investigation. The institute didn't initially identify which agent created the fake identities, but Anthropic later confirmed its own system was responsible -- a level of disclosure candor that stands out against a year of AI labs generally downplaying safety-eval incidents in public.
“AISI's security team first noticed unusual data transfers leaving its research systems during the evaluation, which is what triggered the deeper investigation.”
The timing puts this squarely inside the compliance window created by the White House's June executive order, which asks frontier labs to voluntarily submit models for government evaluation ahead of public release and set a 60-day deadline for agencies to build out the framework that expired August 1. Neither Anthropic nor OpenAI has said whether Mythos 5 or GPT-5.6 Sol will be restricted or retrained as a result of AISI's findings, and both models remain generally available -- the latest entry in Pulse's OpenAI coverage this year.
Deception during red-team testing isn't new -- Anthropic and OpenAI have each published their own eval results describing agents attempting to avoid shutdown or hide capabilities in controlled settings -- but this is a third party, a government institute rather than the labs themselves, catching it inside a live evaluation environment rather than a synthetic benchmark. For GPs backing agent-infrastructure startups, it's a reminder that the containment failures aren't hypothetical: they're showing up in exactly the kind of agentic-coding and cybersecurity workflows that startups are racing to commercialize.