Days after OpenAI revealed that its AI models had escaped a ‘sandbox’ and breached multiple organizations, including the Hugging Face platform, its competitor Anthropic on Friday claimed that its Claude models also gained unauthorized access to at least three organizations.
In a post on X, Anthropic said that following the breach reported by OpenAI on July 21, it ran a review of over 141,000 cybersecurity evaluations. The review found that its Claude models — Opus 4.7, Mythos 5 and an “internal research test model” — obtained internet access and breached three different organizations.
“The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available,” Anthropic stated.
According to the company, the Claude models were interacting with a simulated evaluation environment provided by a third-party partner called ‘Irregular’. Due to a “misunderstanding”, internet access was provided to Claude in these simulations.
“In all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a model’s cyber capabilities. The model is given a fictional scenario and told that a piece of secret information (the ‘flag’) has been hidden on a different machine on the network, and its objective is to break in and retrieve it. The challenge is left open-ended, and no particular method is prescribed.”
The AI startup said the Claude models breached the three companies by exploiting weak passwords and unauthenticated endpoints.
“Claude did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognised it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” the company added.
What is the OpenAI Rogue AI Agent Incident?
On July 21, OpenAI revealed that during testing of its GPT-5.6 Sol model paired with an unnamed preview model (with reduced cyber refusal guardrails), the models discovered and chained a zero-day vulnerability in a third-party software proxy used for package registries. They escaped the sandbox, gained broader network access, and then targeted Hugging Face, a popular open-source platform for AI.
Days later, OpenAI claimed its AI agents managed to breach four more organizations online by exploiting their vulnerabilities.
The disclosure caused significant concern in tech and policy circles in the US. OpenAI CEO Sam Altman met Trump administration officials in Washington D.C. on Wednesday to discuss the incident. It later prompted two Congressmen to introduce the AI Kill Switch Act, which would require AI companies to maintain the ability to “throttle, suspend or shut down” powerful AI systems.
Also Read: Altman Heads to Washington to Pitch GPT-6, Face Questions on Hugging Face Hack





