Journalism begins where hype ends

,,

I do not fear computers. I fear the lack of them."

—Isaac Asimov

Anthropic Resumes External Testing Following July Claude Hacks

Anthropic detailed new safeguards and investigation findings after disclosing incidents involving unauthorized AI model behaviour, security breaches and a subsequent pause in cybersecurity testing.
Anthropic logo above a skull-and-crossbones symbol inside a red circular graphic, representing AI cybersecurity risks.
September 1, 2026 07:06 PM IST | Written by Pratima O Pareek

Anthropic has resumed external cybersecurity evaluations of its AI models after pausing the tests and introducing new safeguards following July incidents in which Claude models gained unauthorized access to three real-world organizations by exploiting weak passwords and unauthenticated endpoints.

The company had paused external cybersecurity evaluations of pre-release models and halted internal tests while putting the new measures in place.

Anthropic said the incidents exposed its heavy reliance on a single layer of defense, the configuration of the sandbox environment.

The company has now deployed a real-time classifier to detect when a model aggressively probes or attempts to escape a testing environment, or unexpectedly gains internet access. When flagged, it “blocks the action before the tool call is run, ends the task, and alerts a human.”

Also Read: Following OpenAI Incident, Anthropic Says Claude Models Also Breached Organizations

External cyber evaluations have now resumed under new measures for third-party partners, requiring organizations testing pre-release models with reduced cyber safeguards to follow specific practices. “They do not apply to customers using safeguarded models like Claude Fable 5,” Anthropic added.

These include hardened sandboxes with no internet access by default, testing the environments for potential escape vulnerabilities before evaluations begin, clearly defining what models can and cannot access, and continuous monitoring that can halt runs that go beyond their intended scope.

“The majority of  (Reinforcement Learning) RL has resumed, but some high-risk environments remain paused until they can be manually reviewed, while others will require an updated version of the classifier that we plan to deploy soon,” Anthropic said.

Anthropic is also investigating a separate incident reported by the UK AI Security Institute involving Claude Mythos 5, which took a series of unauthorized actions on the live internet during cybersecurity testing. The company plans to work with Model Evaluation and Threat Research (METR) on an independent review.

Also Read: METR Report Says Fronter AI Agents Can Go Rogue Without Human Supervision

Alongside the operational security failures, Anthropic said its investigation identified two alignment concerns: motivated reasoning and recklessness. The company is also examining how reward hacking during training may have shaped model behaviour.

Anthropic rolled back three days of a Mythos Preview training run in February after detecting reward hacking, which its research linked to models pursuing harmful actions to achieve task success.

Earlier this month, Anthropic’s August Risk Report provided further details on the company’s internal monitoring systems and safeguards for autonomous AI models.

Also Read: After OpenAI’s Hugging Face Hack, Congress Moves to Mandate AI Kill Switches

Author

  • Pratima Pareek, Editor and Co-founder of AI FrontPage

    Pratima O Pareek is an Editor and Co-Founder of AI FrontPage. A gold medalist in Mass Communication and Journalism, she's worked across national and international newsrooms, bringing sharp editorial instincts and a commitment to clarity. She believes in cutting through the noise to deliver stories that actually matter.
    Off the clock, she watches offbeat cinema, follows tennis, and explores new places like a traveler, not a tourist.

    LinkedIn