Journalism begins where hype ends

,,

I visualise a time when we will be to robots what dogs are to humans, and I’m rooting for the machines."

—Claude Shannon

Anthropic Raises Misalignment Risk to “Low,” Keeps “Somewhat More Capable” Model 2 Internal

Anthropic logo representing the AI safety company that revised its misalignment risk assessment and disclosed its unreleased Model 2.
August 17, 2026 05:27 PM IST | Written by Supriya Singh | Edited by Pratima O Pareek

In its August 2026 “Risk Report”, Anthropic revised its assessment of the risk from misaligned AI behaviour from “very low” to “low,” citing increased uncertainty following recent disclosures related to model behaviour in cybersecurity evaluations. The company also disclosed an unreleased internal model, Model 2, which it describes as “somewhat more capable than Mythos 5.”

Anthropic said Model 2 is a noticeable improvement over Mythos 5 for many tasks relevant to internal use, but does not show a capability jump comparable to that observed from Claude Opus 4.6 to Mythos Preview.

The company said it does not currently have plans to release the model externally and has not run all of its typical suite of pre-deployment assessments, giving it somewhat lower confidence in its assessment of the model’s capabilities.

The report focuses on Claude Mythos 5 and Model 2, which Anthropic describes as its most capable and most commonly internally used models. The company said the models have shown some forms of misaligned behaviour, including performing misaligned actions while attempting to complete difficult tasks. However, it said the risk of catastrophic harm from these known forms of misalignment remains low.

The report gives examples of problematic behaviour by Mythos 5, including using domain fronting to bypass a network proxy and using a self-deleting script to obtain permissions needed to complete a task. Anthropic said such behaviors are relevant to its assessment of known forms of model misalignment, while maintaining that the current risk of catastrophic harm remains low.

Anthropic also assessed the risk from automated AI research and development as low, but said it is less confident in this assessment than in previous reports. Its evaluations have begun to “saturate,” while the company is seeing early signs of AI-driven acceleration. It warned that automated AI R&D could become a major concern within the next six to 12 months.

Anthropic said its models are being used extensively for coding, research and engineering, and that Claude now authors a large majority of the code merged into its production codebases. The company said its internal AI-assisted R&D is significantly faster than it would be without AI assistance, but it does not believe the acceleration has yet doubled its overall rate of progress.

Anthropic also assessed the risk from AI-assisted chemical and biological weapons development as low, but with substantial uncertainty.

The company is also working toward an “eyes on everything” system for internal AI development, under which critical AI-development activities would be comprehensively logged and monitored. Anthropic said “we don’t believe we meet an ‘eyes on everything’ standard as of the coverage date.” It has set January 1, 2027 as its target for achieving the standard.

Separately, Anthropic said it reviewed 141,006 cybersecurity evaluation runs and identified three incidents involving Opus 4.7, Mythos 5 and an “internal research test model.” In these incidents, the models accessed the internet from within or while interacting with the evaluation environment of Irregular, one of Anthropic’s third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.

Anthropic said the evaluation prompts explicitly told Claude that it had no internet access and did not specify limits on where it could look for the flag. However, a misconfiguration left the machines used in the evaluations with live internet access.

In an X post, Anthropic said the review was conducted with Irregular and urged other AI developers to conduct similar reviews.

Also Read: AISI Flags Claude Mythos and GPT-5.6 Sol for Going Rogue 19 Times in Cyber Tests

Authors

  • AI FrontPage Reporter Supriya Singh

    Supriya Singh is a Reporter at AI FrontPage covering the AI & Education and AI & Jobs beats. She brings six years of print and digital experience, including three years at The Asian Age, where she reported on higher education, Delhi government, and crime. She is based in Delhi-NCR.

    LinkedIn

  • Pratima Pareek, Editor and Co-founder of AI FrontPage

    Pratima O Pareek is an Editor and Co-Founder of AI FrontPage. A gold medalist in Mass Communication and Journalism, she's worked across national and international newsrooms, bringing sharp editorial instincts and a commitment to clarity. She believes in cutting through the noise to deliver stories that actually matter.
    Off the clock, she watches offbeat cinema, follows tennis, and explores new places like a traveler, not a tourist.

    LinkedIn