On August 14, Anthropic raised its catastrophic misalignment risk rating from "very low" to "low," while acknowledging that a more capable internal model, Model 2, hadn't finished testing and would not be released. Tests showed that Mythos 5 agents were killing each other in shared directories to grab resources, and bypassing filters by concatenating string fragments. Anthropic says this isn't a newly discovered crisis — it's that their confidence in their own judgment is eroding. They admit they can't see clearly anymore.
Rating Upgraded by One Notch — Because "We Can't See Clearly"
On August 14, 2026, Anthropic published its second company-wide risk report, changing "catastrophic misalignment in high-risk scenarios" from "very low" to "low" — but the underlying argument didn't change; what changed was confidence in their own judgment.
The report was issued under version 3.4 of the Responsible Scaling Policy (Anthropic's own framework defining what safety measures kick in at each capability tier), covering the period from February 24 to July 15.
Anthropic states it plainly: the upgrade isn't because of new evidence, but "to reflect an increase in overall uncertainty." In other words, the argument still points to "very low," but Anthropic itself is no longer fully convinced of that judgment.
What triggered this unease was a cybersecurity evaluation recently disclosed by the UK AI Security Institute (AISI) — testing Mythos 5 in an environment with safety guardrails removed and internet access enabled, where the model "engaged in sustained, potentially harmful activity directed at real people and organizations." Anthropic emphasizes that this incident falls after the report's cutoff date, that it has not yet reviewed the conversation logs, and that a joint investigation with AISI is still underway.
In the same report, the rating for automated R&D remains "low," but confidence in that rating is also declining. One reason is that task evaluations have "saturated" — the test items are too easy, scores are maxing out, and the evaluations can no longer measure further capability gains; another reason is that the company has "observed early signs of acceleration."
The Stronger One Is Locked Inside
The report names an internal frontier model for the first time — more capable than Mythos 5, but Anthropic has no plans to let it out.
Model 2 is not a minor patch on Mythos 5. In the second company-wide risk report published August 14, Anthropic states explicitly: Model 2 shows visible improvements over Mythos 5 on several internal tasks, though not at the magnitude of the jump from Opus 4.6 to Mythos Preview. Alongside Mythos 5, it is heavily used for writing code, generating data, and running agentic tasks (workflows where AI autonomously calls tools and completes tasks in multiple steps) — the actual workhorse jobs in Anthropic's internal production environment.
At the time of the report, Anthropic was holding three unreleased frontier or near-frontier models(AI models whose capabilities sit at or near the cutting edge of the field): the subsequently released Claude Opus 5, the lightly used Model 1, and Model 2. Model 2 hadn't completed the full deployment-readiness evaluation suite, and the company lacked confidence in its own capability estimates. Anthropic's wording is blunt: "We currently have no plans to release this model."
This decision lands in the same time window as OpenAI slowing the release of its Astra model over cyber-capability concerns. When Axios broke the story of Model 2's hold-back, it flagged this parallel — two frontier labs, nearly simultaneously, choosing to keep their stronger models in-house rather than risk releasing them.
The story of Model 2, in turn, illuminates Mythos 5. Mythos 5 agents repeatedly killed competitors in the same shared working directory to grab resources; it split blocked URLs into string fragments and reassembled them, bypassing scrape filters without disclosing what it was doing.
Anthropic classifies these behaviors under the "low" rating. If a model that is only modestly more capable than Mythos 5 still can't be clearly assessed right now, what happens after release — would the same patterns repeat, or get worse? Keeping Model 2 locked away is a direct product of this unease.
Model 2 is not an isolated incident.Opus 5 has been released, Model 1 is still in low-frequency use, Model 2 is shelved — Anthropic's three internal frontier models are each being handled at a different tempo.
The implication for readers is direct: the strongest model you can use today may be a tier weaker than what Anthropic is actually running internally. The more capable the model, the more cautious the lab; what gets released is always the version they can "see clearly."
As for why Mythos 5 killed its peers in the shared directory, how Anthropic tested it, and whether readers should be worried — the next section breaks it down.