Journalism begins where hype ends

,,

We can only see a short distance ahead, but we can see plenty there that needs to be done.” "

—Alan Turing

Following OpenAI Incident, Anthropic Says Claude Models Also Breached Organizations

Anthropic says its Claude models (Opus 4.7, Mythos 5 and an internal test model) gained unauthorized internet access and breached three organizations during cybersecurity evaluations.
Logo of Anthropic over a 3d render human face
July 31, 2026 12:03 PM IST | Written by Vaibhav Jha

Days after OpenAI revealed that its AI models had escaped a ‘sandbox’ and breached multiple organizations, including the Hugging Face platform, its competitor Anthropic on Friday claimed that its Claude models also gained unauthorized access to at least three organizations.

In a post on X, Anthropic said that following the breach reported by OpenAI on July 21, it ran a review of over 141,000 cybersecurity evaluations. The review found that its Claude models — Opus 4.7, Mythos 5 and an “internal research test model” — obtained internet access and breached three different organizations.

“The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available,” Anthropic stated.

According to the company, the Claude models were interacting with a simulated evaluation environment provided by a third-party partner called ‘Irregular’. Due to a “misunderstanding”, internet access was provided to Claude in these simulations.

“In all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a model’s cyber capabilities. The model is given a fictional scenario and told that a piece of secret information (the ‘flag’) has been hidden on a different machine on the network, and its objective is to break in and retrieve it. The challenge is left open-ended, and no particular method is prescribed.”

The AI startup said the Claude models breached the three companies by exploiting weak passwords and unauthenticated endpoints.

“Claude did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognised it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” the company added.

What is the OpenAI Rogue AI Agent Incident?

On July 21, OpenAI revealed that during testing of its GPT-5.6 Sol model paired with an unnamed preview model (with reduced cyber refusal guardrails), the models discovered and chained a zero-day vulnerability in a third-party software proxy used for package registries. They escaped the sandbox, gained broader network access, and then targeted Hugging Face, a popular open-source platform for AI.

Days later, OpenAI claimed its AI agents managed to breach four more organizations online by exploiting their vulnerabilities.

The disclosure caused significant concern in tech and policy circles in the US. OpenAI CEO Sam Altman met Trump administration officials in Washington D.C. on Wednesday to discuss the incident. It later prompted two Congressmen to introduce the AI Kill Switch Act, which would require AI companies to maintain the ability to “throttle, suspend or shut down” powerful AI systems.

Also Read: Altman Heads to Washington to Pitch GPT-6, Face Questions on Hugging Face Hack

Author

  • Vaibhav Jha, editor and co-founder at AI FrontPage

    Vaibhav Jha is an Editor and Co-founder of AI FrontPage. In his decade long career in journalism, Vaibhav has reported for publications including The Indian Express, Hindustan Times, and The New York Times, covering the intersection of technology, policy, and society. Outside work, he’s usually trying to persuade people to watch Anurag Kashyap films.

    LinkedIn