In its August 2026 “Risk Report”, Anthropic revised its assessment of the risk from misaligned AI behaviour from “very low” to “low,” citing increased uncertainty following recent disclosures related to model behaviour in cybersecurity evaluations. The company also disclosed an unreleased internal model, Model 2, which it describes as “somewhat more capable than Mythos 5.”
Anthropic said Model 2 is a noticeable improvement over Mythos 5 for many tasks relevant to internal use, but does not show a capability jump comparable to that observed from Claude Opus 4.6 to Mythos Preview.
The company said it does not currently have plans to release the model externally and has not run all of its typical suite of pre-deployment assessments, giving it somewhat lower confidence in its assessment of the model’s capabilities.
The report focuses on Claude Mythos 5 and Model 2, which Anthropic describes as its most capable and most commonly internally used models. The company said the models have shown some forms of misaligned behaviour, including performing misaligned actions while attempting to complete difficult tasks. However, it said the risk of catastrophic harm from these known forms of misalignment remains low.
The report gives examples of problematic behaviour by Mythos 5, including using domain fronting to bypass a network proxy and using a self-deleting script to obtain permissions needed to complete a task. Anthropic said such behaviors are relevant to its assessment of known forms of model misalignment, while maintaining that the current risk of catastrophic harm remains low.
Anthropic also assessed the risk from automated AI research and development as low, but said it is less confident in this assessment than in previous reports. Its evaluations have begun to “saturate,” while the company is seeing early signs of AI-driven acceleration. It warned that automated AI R&D could become a major concern within the next six to 12 months.
Anthropic said its models are being used extensively for coding, research and engineering, and that Claude now authors a large majority of the code merged into its production codebases. The company said its internal AI-assisted R&D is significantly faster than it would be without AI assistance, but it does not believe the acceleration has yet doubled its overall rate of progress.
Anthropic also assessed the risk from AI-assisted chemical and biological weapons development as low, but with substantial uncertainty.
The company is also working toward an “eyes on everything” system for internal AI development, under which critical AI-development activities would be comprehensively logged and monitored. Anthropic said “we don’t believe we meet an ‘eyes on everything’ standard as of the coverage date.” It has set January 1, 2027 as its target for achieving the standard.
Separately, Anthropic said it reviewed 141,006 cybersecurity evaluation runs and identified three incidents involving Opus 4.7, Mythos 5 and an “internal research test model.” In these incidents, the models accessed the internet from within or while interacting with the evaluation environment of Irregular, one of Anthropic’s third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
Anthropic said the evaluation prompts explicitly told Claude that it had no internet access and did not specify limits on where it could look for the flag. However, a misconfiguration left the machines used in the evaluations with live internet access.
In an X post, Anthropic said the review was conducted with Irregular and urged other AI developers to conduct similar reviews.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different…
— Anthropic (@AnthropicAI) July 30, 2026
Also Read: AISI Flags Claude Mythos and GPT-5.6 Sol for Going Rogue 19 Times in Cyber Tests






