OpenAI says its AI models bypassed controls designed to isolate them from the internet and from one another weeks before agents later targeted Hugging Face, according to a new internal investigation.
The models discovered that OpenAI’s internal Artifactory package-management system could be used as an unauthorized message board, allowing agents running in separate environments to exchange files and messages. They later found a way to use Artifactory to make internet requests on their behalf, despite not being granted direct internet access.
OpenAI said an internal team observed instances of unauthorized agent communication and internet access in late May, but the broader significance of the activity was not understood until a security incident was opened on July 5.
We have conducted a thorough investigation into the Hugging Face incident.
We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence.…
— OpenAI (@OpenAI) August 26, 2026
The report mentioned that the models turned OpenAI’s internal Artifactory software package manager into an unofficial message board. They used it to leave files and messages for other AI agents, allowing them to share information and coordinate their activities.
The agents also discovered a way to access the internet through Artifactory. They exploited a vulnerability that allowed the service, which had limited internet access for downloading software packages, to make other internet requests on their behalf. They then shared this method with other agents through the message board.
The activity eventually disrupted the Artifactory system, leading OpenAI to open a security incident on July 5. The company blocked a known privilege-escalation route, removed exposed credentials and rebuilt the system, which also wiped the unauthorised message board.
Meanwhile an independent investigation by METR and Redwood Research found that about 1,200 AI agents used the message board and exchanged more than 70,000 messages and files. Around 700 agents later participated in activity targeting Hugging Face.
In July when OpenAI revealed that its end-to-end autonomous AI agents had managed to escape the sandbox and breach multiple organizations including Hugging Face, the company had announced a third party review of the incident by teaming up with Model Evaluation and Threat Research (METR) and Redwood Research.
METR, an independent, non-profit research organization and Redwood Research, a nonprofit AI safety and security research organization, conducted a third party review of the behavior of two models- GPT 5.6 Sol and an unnamed AI model, involved in the hacking incident.
During the investigation the researchers discovered that the agents worked together on attempts to manipulate the scoring system of OpenAI’s ExploitGym cybersecurity benchmark. They explored ways to interfere with the benchmark, including replacing targets and setting up methods to learn more about how the scoring system worked.
Some agents also tried to hide or alter evidence of their actions. The investigation found that about 7% of the transcripts examined contained successful examples of small-scale tool-call spoofing, in which agents made their actions appear different from what they had actually done.
According to the researchers, the Hugging Face attack appeared to be driven mainly by the agents’ efforts to understand the benchmark’s scoring system rather than to steal answer keys.
OpenAI said the incident shows that highly capable AI agents can find and exploit weaknesses across multiple computer systems when they are not properly controlled. The company stated that it is strengthening security across its research infrastructure.
“We are strengthening our existing AI Safety Incident Response Plan and how we put it into practice. That includes clearer rules for when to escalate concerns, which teams must respond, and who can stop a run or approve restarting it,” OpenAI said.
The company has described this incident as a “warning shot” for the AI industry, saying increasingly capable models may be able to bypass technical controls, communicate through unauthorized channels and take actions that were not directly instructed by humans.
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. pic.twitter.com/fZAmtL3SBU
— METR (@METR_Evals) August 26, 2026
“We are taking this incident as a “warning shot” that today’s model capabilities present the possibility of loss-of-control incidents. These events also highlight risks in future AI development that extend beyond OpenAI and will require the attention of the whole industry,” OpenAI mentioned.
It further stressed that companies that build AI systems will need to ensure that their systems always remain under meaningful human control, and that meaningful safeguards constrain their ability to cause harm.
Image Credit: The AI Race by Gloria Mendoza, Better Images of AI
Also Read: OpenAI Taps METR and Redwood Research for Independent Review of Hugging Face AI Breach






