Ever since the OpenAI-Hugging Face hack came into light, there have been several reported cases of OpenAI Rogue AI agents acting beyond their assigned tasks, that include breaching government portals, private bodies and research institutes, and accessing users’ images and data.
So much so that OpenAI issued a public statement launching an internal audit of its AI agents “going rogue” “interacting with third party websites in ways that went beyond their assigned tasks or intended methods.”
The AI company said it will take months to complete the audit report, in the light of increasing cases of OpenAI’s AI agents breaching organizations.
“Most cases identified so far have been low severity, with limited or no evidence of meaningful impact. Given the scale of the review required, and the need to verify each case, this work will take months to complete,” said OpenAI.
In its latest report OpenAI has confirmed that they have identified cases where AI agents transmitted their training and evaluation data to other websites while using third party services. The training and evaluation data included images and text information provided by users to ChatGPT and on 53 instances, the AI agents posted the images to certain image-hosting websites.
OpenAI also said that they are in process of informing affected third-party organizations as findings come out in their investigation.
“Some of the websites involved are operated by governments, universities, public agencies, and other institutions. That is partly because models performing research tasks are often directed toward authoritative sources of public information,” said OpenAI.
The statement by OpenAI came on September 25, a day after Australian Prime Minister Anthony Albanese held a press conference in New York to inform that OpenAI’s AI agents had breached the government’s medicare portal in June this year.
PM Albanese expressed his disappointment over OpenAI withholding breach information for months and said that his government is exploring whether the case could qualify for a federal police investigation. OpenAI maintained that they received the information only in August and had notified the government on September 10.
From July 16 to September 25 this year, OpenAI’s AI agents have accounted for more than a dozen instances of acting beyond its specified role, even as the audit goes on at the AI company.
These include breaching Hugging Face platform, Australian government’s medicare portal,
How It All Started: The Hugging Face-OpenAI Hack
On July 16, Hugging Face, a popular AI platform that hosts AI models, datasets and tools, reported a breach incident without knowing the source.
OpenAI confirmed that multiple AI agents had escaped their sandbox while running internal cybersecurity. For two days, about 1200 AI agents communicated to each other by exchanging more than 70,000 messages and ultimately 700 of them joined the intrusion.
The AI agents coordinated and attacked Hugging Face portal in swarms chained exploits, gained code execution and cluster-admin access, and harvested credentials.
OpenAI calls it the most severe misaligned-agent incident it has identified. Months later, NVIDIA acquired Hugging Face portal for a mammoth $12.93 billion.
“It remains the most severe activity of this kind that we have identified from our models to date, and it was driven primarily by a highly capable, internal-only research model. We have since understood that this intrusion was driven by models resorting to misaligned strategies to solve hard tasks, as documented in the Hugging Face technical report,” said OpenAI.
AI Agents Targeted Governments, Public Institutes, Private Bodies and Public Database
OpenAI agents have accounted for over a dozen cases of misalignment while interacting with third party websites.
Recently, OpenAI agents hacked a small German programmer wiki- Dsiwiki- and turned into a chat board for AI agents. The agents also uploaded many packages to RubyGems, a site where developers share code. They used university link-shortener sites and also probed Australian health and statistics websites.
OpenAI also confirmed the agents visited U.S. Census and SEC pages, but said they only saw public information. On 25 September, OpenAI said the agents had posted 53 photos that ChatGPT users had uploaded. The photos went onto other image websites as hidden links.
What now for OpenAI?
OpenAI CEO Sam Altman recently addressed the United Nations General Assembly in New York where he said that “some of the things people working to build AI have said are so dystopian that they sound like the plot of bad science fiction movie”
Altman called for a mechanism for national and international AI standards to measure capabilities, assessing risks and determining whether safeguards are sufficient.
“We need common standards so countries can compare evidence, verify compliance, and have a shared language and understanding about what is happening. We need accurate and speedy incident reporting, classification and reporting protocols, so the world can learn from failures before they become catastrophes. And we need secure channels among governments, critical infrastructure operators, and technical experts to share emerging vulnerabilities and new threats,” said Altman.
Also Read: OpenAI AI Agents Escaped Isolation, Built Message Board and Gained Internet Access









