Analysis
Anthropic published its second company-wide Risk Report on August 14, disclosing the existence of an unreleased internal model, Model 2, that outperforms its current public flagship Mythos 5 -- scoring 62.8% on Anthropic's internal CoBench benchmark versus 50.3% for Mythos 5, according to SiliconANGLE. Model 2 sits in the same 'Mythos' capability tier as the released model, but Anthropic says it hasn't completed its full predeployment assessment suite, meaning the company has lower confidence in its capability and safety claims than it does for systems it actually ships -- and has no current plans to release it.
The more consequential change in the report is a shift in Anthropic's own risk rating: the company raised its assessment of catastrophic harm from misalignment in high-stakes settings from 'very low,' its rating in the first Risk Report published in February 2026, to 'low.' Critically, per TechTimes, the change wasn't triggered by a model failing a new safety test -- it followed increased uncertainty from recent cybersecurity-evaluation incident disclosures. The report explicitly states Model 2's internal review surfaced no new or more concerning misalignment behavior beyond what was already documented for Mythos 5.
“The report explicitly states Model 2's internal review surfaced no new or more concerning misalignment behavior beyond what was already documented for Mythos 5.”
The disclosure comes as Anthropic's business fundamentals are accelerating -- the company reported $11.5 billion in Q2 revenue this week, and recently hired Mariano-Florentino Cuéllar as Chief Global Affairs Officer, signaling growing attention to the regulatory and safety-policy side of the business alongside commercial growth. Anthropic has positioned its safety-first branding as a competitive differentiator against OpenAI and Google DeepMind, both of which release capability benchmarks more aggressively and have faced more public criticism over safety-evaluation transparency.
What the report doesn't resolve is the tension at its core: a lab voluntarily raising its own risk rating, in a public document, while simultaneously disclosing a more capable unreleased model, is a genuinely unusual piece of self-regulation in an industry where most labs disclose benchmarks that make them look better, not more cautious. It's also a reminder that 'low' risk is still a real, non-zero rating from the company's own internal framework -- not a clean bill of health.
For investors evaluating AI-safety-adjacent bets, the report is a data point on how seriously the frontier labs are treating self-assessment as regulatory scrutiny increases, rather than a signal about near-term commercial risk. Anthropic's revenue trajectory hasn't been affected by the disclosure, and the company continues to prepare for a public listing later this year.