Journalism begins where hype ends

,,

We can only see a short distance ahead, but we can see plenty there that needs to be done.” "

—Alan Turing

OpenAI Taps METR and Redwood Research for Independent Review of Hugging Face AI Breach

Days after OpenAI revealed that its autonomous AI agents escaped a testing sandbox and breached Hugging Face (and other organizations), the company has engaged independent nonprofits METR and Redwood Research to review the model behavior.
Logos of METR, OpenAI and Redwood Research from left to right.
July 31, 2026 02:30 PM IST | Written by Neelam Sharma | Edited by Vaibhav Jha

Days after OpenAI revealed that its end-to-end autonomous AI agents had managed to escape the sandbox and breach multiple organizations including Hugging Face, the company announced a third party review of the incident by teaming up with Model Evaluation and Threat Research (METR) and Redwood Research.

METR, an independent, non-profit research organization and Redwood Research, a nonprofit AI safety and security research organization, will be conducting a third party review of the behavior of two models- GPT 5.0 Sol and an unnamed AI model, involved in the hacking incident.

“We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions,” informed METR on a post on X on Wednesday.

 

On July 21, OpenAI revealed that during testing of its GPT-5.6 Sol model paired with an unnamed preview model, with reduced cyber refusal guardrails, the models discovered and chained a zero-day vulnerability in a third party software proxy used for package registries, escaped the sandbox, gained broader network access, and then target Hugging Face, a popular open-source platform for AI.

Days later, OpenAI found that its AI agents had breached at least four more organizations apart from Hugging Face by exploiting vulnerabilities. One of the four was New York-based Modal Labs whose Chief Technology Officer Akshat Bubna spoke to Reuters, explaining that the mishap was the result of a lack of security in the code shared by one of the clientele of the firm.

The revelations resulted in a major controversy with demands for increased safeguards around AI and advocacy for open-source AI models along with introduction of an AI Kill Switch Act by two congressmen.

On Friday, Anthropic announced that following OpenAI’s announcement on the hugging face hacking incident, it ran a review of over 141,000 simulated cybersecurity evaluations of its Claude AI models and found that they breached at least three different organizations.

Also Read: Altman Heads to Washington to Pitch GPT-6, Face Questions on Hugging Face Hack

Authors

  • Neelam Sharma, reporter at AI FrontPage

    Neelam Sharma is a passionate storyteller, and journalist with over a decade of experience across leading Indian media houses.
    Known for her calm presence on screen and powerful storytelling off it, Neelam brings a rare blend of credibility, creativity, and empathy to journalism. Her strength lies in ground reporting and research-driven narratives that connect with the heart of the audience. Whether covering social issues, human-interest features, or breaking news, she combines factual depth with a human touch—making every story not just informative.

    LinkedIn

  • Vaibhav Jha, editor and co-founder at AI FrontPage

    Vaibhav Jha is an Editor and Co-founder of AI FrontPage. In his decade long career in journalism, Vaibhav has reported for publications including The Indian Express, Hindustan Times, and The New York Times, covering the intersection of technology, policy, and society. Outside work, he’s usually trying to persuade people to watch Anurag Kashyap films.

    LinkedIn