In its August 2026 “Risk Report”, Anthropic revised its assessment of the risk from misaligned AI behaviour from “very low” to “low,” citing increased uncertainty following recent disclosures related to model behaviour in cybersecurity evaluations. The company also disclosed an unreleased internal model, Model 2, which it describes as “somewhat more capable than Mythos 5.”… Continue reading Anthropic Raises Misalignment Risk to “Low,” Keeps “Somewhat More Capable” Model 2 Internal
Anthropic Raises Misalignment Risk to “Low,” Keeps “Somewhat More Capable” Model 2 Internal