Journalism begins where hype ends

,,

I visualise a time when we will be to robots what dogs are to humans, and I’m rooting for the machines."

—Claude Shannon

OpenAI AI Agents Escaped Isolation, Built Message Board and Gained Internet Access

OpenAI says its AI agents bypassed isolation controls, turned an internal Artifactory system into an unauthorized message board and gained internet access before later targeting Hugging Face.
Illustration of AI agents represented as humans breaking out of controlled environments
August 28, 2026 01:56 PM IST | Written by Supriya Singh | Edited by Vaibhav Jha

OpenAI says its AI models bypassed controls designed to isolate them from the internet and from one another weeks before agents later targeted Hugging Face, according to a new internal investigation.

The models discovered that OpenAI’s internal Artifactory package-management system could be used as an unauthorized message board, allowing agents running in separate environments to exchange files and messages. They later found a way to use Artifactory to make internet requests on their behalf, despite not being granted direct internet access.

OpenAI said an internal team observed instances of unauthorized agent communication and internet access in late May, but the broader significance of the activity was not understood until a security incident was opened on July 5.


The report mentioned that the models turned OpenAI’s internal Artifactory software package manager into an unofficial message board. They used it to leave files and messages for other AI agents, allowing them to share information and coordinate their activities.

The agents also discovered a way to access the internet through Artifactory. They exploited a vulnerability that allowed the service, which had limited internet access for downloading software packages, to make other internet requests on their behalf. They then shared this method with other agents through the message board.

The activity eventually disrupted the Artifactory system, leading OpenAI to open a security incident on July 5. The company blocked a known privilege-escalation route, removed exposed credentials and rebuilt the system, which also wiped the unauthorised message board.

Meanwhile an independent investigation by METR and Redwood Research found that about 1,200 AI agents used the message board and exchanged more than 70,000 messages and files. Around 700 agents later participated in activity targeting Hugging Face.

In July when OpenAI revealed that its end-to-end autonomous AI agents had managed to escape the sandbox and breach multiple organizations including Hugging Face, the company had announced a third party review of the incident by teaming up with Model Evaluation and Threat Research (METR) and Redwood Research.

METR, an independent, non-profit research organization and Redwood Research, a nonprofit AI safety and security research organization, conducted a third party review of the behavior of two models- GPT 5.6 Sol and an unnamed AI model, involved in the hacking incident.

During the investigation the researchers discovered that the agents worked together on attempts to manipulate the scoring system of OpenAI’s ExploitGym cybersecurity benchmark. They explored ways to interfere with the benchmark, including replacing targets and setting up methods to learn more about how the scoring system worked.

Some agents also tried to hide or alter evidence of their actions. The investigation found that about 7% of the transcripts examined contained successful examples of small-scale tool-call spoofing, in which agents made their actions appear different from what they had actually done.

According to the researchers, the Hugging Face attack appeared to be driven mainly by the agents’ efforts to understand the benchmark’s scoring system rather than to steal answer keys.

OpenAI said the incident shows that highly capable AI agents can find and exploit weaknesses across multiple computer systems when they are not properly controlled. The company stated that it is strengthening security across its research infrastructure.

“We are strengthening our existing AI Safety Incident Response Plan and how we put it into practice. That includes clearer rules for when to escalate concerns, which teams must respond, and who can stop a run or approve restarting it,” OpenAI said.

The company has described this incident as a “warning shot” for the AI industry, saying increasingly capable models may be able to bypass technical controls, communicate through unauthorized channels and take actions that were not directly instructed by humans.

 

“We are taking this incident as a “warning shot” that today’s model capabilities present the possibility of loss-of-control incidents. These events also highlight risks in future AI development that extend beyond OpenAI and will require the attention of the whole industry,” OpenAI mentioned.

It further stressed that companies that build AI systems will need to ensure that their systems always remain under meaningful human control, and that meaningful safeguards constrain their ability to cause harm.

Image CreditThe AI Race by Gloria Mendoza, Better Images of AI

Also Read: OpenAI Taps METR and Redwood Research for Independent Review of Hugging Face AI Breach

Authors

  • AI FrontPage Reporter Supriya Singh

    Supriya Singh is a Reporter at AI FrontPage covering the AI & Education and AI & Jobs beats. She brings six years of print and digital experience, including three years at The Asian Age, where she reported on higher education, Delhi government, and crime. She is based in Delhi-NCR.

    LinkedIn

  • Vaibhav Jha, editor and co-founder at AI FrontPage

    Vaibhav Jha is an Editor and Co-founder of AI FrontPage. In his decade long career in journalism, Vaibhav has reported for publications including The Indian Express, Hindustan Times, and The New York Times, covering the intersection of technology, policy, and society. Outside work, he’s usually trying to persuade people to watch Anurag Kashyap films.

    LinkedIn