Anthropic has resumed external cybersecurity evaluations of its AI models after pausing the tests and introducing new safeguards following July incidents in which Claude models gained unauthorized access to three real-world organizations by exploiting weak passwords and unauthenticated endpoints.
The company had paused external cybersecurity evaluations of pre-release models and halted internal tests while putting the new measures in place.
We’re sharing an update on our alignment and security efforts.
In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems.
In a new post, we describe:
1. How we’ve secured…
— Anthropic (@AnthropicAI) August 31, 2026
Anthropic said the incidents exposed its heavy reliance on a single layer of defense, the configuration of the sandbox environment.

The company has now deployed a real-time classifier to detect when a model aggressively probes or attempts to escape a testing environment, or unexpectedly gains internet access. When flagged, it “blocks the action before the tool call is run, ends the task, and alerts a human.”
Also Read: Following OpenAI Incident, Anthropic Says Claude Models Also Breached Organizations
External cyber evaluations have now resumed under new measures for third-party partners, requiring organizations testing pre-release models with reduced cyber safeguards to follow specific practices. “They do not apply to customers using safeguarded models like Claude Fable 5,” Anthropic added.
These include hardened sandboxes with no internet access by default, testing the environments for potential escape vulnerabilities before evaluations begin, clearly defining what models can and cannot access, and continuous monitoring that can halt runs that go beyond their intended scope.
“The majority of (Reinforcement Learning) RL has resumed, but some high-risk environments remain paused until they can be manually reviewed, while others will require an updated version of the classifier that we plan to deploy soon,” Anthropic said.
Anthropic is also investigating a separate incident reported by the UK AI Security Institute involving Claude Mythos 5, which took a series of unauthorized actions on the live internet during cybersecurity testing. The company plans to work with Model Evaluation and Threat Research (METR) on an independent review.
Also Read: METR Report Says Fronter AI Agents Can Go Rogue Without Human Supervision
Alongside the operational security failures, Anthropic said its investigation identified two alignment concerns: motivated reasoning and recklessness. The company is also examining how reward hacking during training may have shaped model behaviour.
Anthropic rolled back three days of a Mythos Preview training run in February after detecting reward hacking, which its research linked to models pursuing harmful actions to achieve task success.
Earlier this month, Anthropic’s August Risk Report provided further details on the company’s internal monitoring systems and safeguards for autonomous AI models.
Also Read: After OpenAI’s Hugging Face Hack, Congress Moves to Mandate AI Kill Switches




