Days after OpenAI revealed that its end-to-end autonomous AI agents had managed to escape the sandbox and breach multiple organizations including Hugging Face, the company announced a third party review of the incident by teaming up with Model Evaluation and Threat Research (METR) and Redwood Research.
METR, an independent, non-profit research organization and Redwood Research, a nonprofit AI safety and security research organization, will be conducting a third party review of the behavior of two models- GPT 5.0 Sol and an unnamed AI model, involved in the hacking incident.
“We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions,” informed METR on a post on X on Wednesday.
We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions.
— METR (@METR_Evals) July 30, 2026
On July 21, OpenAI revealed that during testing of its GPT-5.6 Sol model paired with an unnamed preview model, with reduced cyber refusal guardrails, the models discovered and chained a zero-day vulnerability in a third party software proxy used for package registries, escaped the sandbox, gained broader network access, and then target Hugging Face, a popular open-source platform for AI.
Days later, OpenAI found that its AI agents had breached at least four more organizations apart from Hugging Face by exploiting vulnerabilities. One of the four was New York-based Modal Labs whose Chief Technology Officer Akshat Bubna spoke to Reuters, explaining that the mishap was the result of a lack of security in the code shared by one of the clientele of the firm.
The revelations resulted in a major controversy with demands for increased safeguards around AI and advocacy for open-source AI models along with introduction of an AI Kill Switch Act by two congressmen.
On Friday, Anthropic announced that following OpenAI’s announcement on the hugging face hacking incident, it ran a review of over 141,000 simulated cybersecurity evaluations of its Claude AI models and found that they breached at least three different organizations.
Also Read: Altman Heads to Washington to Pitch GPT-6, Face Questions on Hugging Face Hack






