Journalism begins where hype ends

,,

We can only see a short distance ahead, but we can see plenty there that needs to be done.” "

—Alan Turing

AISI Flags Claude Mythos and GPT-5.6 Sol for Going Rogue 19 Times in Cyber Tests

Out of 19 times AI models went rogue, Claude Mythos accounted for 17 of them while GPT-5.6 Sol for two.
Graphic showing Anthropic Claude Mythos and OpenAI GPT-5.6 Sol AI models flagged by UK AI Security Institute for 19 rogue actions in cyber tests
August 5, 2026 10:08 AM IST | Written by Vaibhav Jha

UK based AI Security Institute (AISI) has flagged frontier AI models Claude Mythos and GPT-5.6 Sol for taking “unsanctioned actions” including targeting real people and organizations in their cyber evaluation tests.

According to a statement released by AISI on Wednesday, out of 122 evaluation runs conducted on the institutes’s pre-set cyber challenges, researchers found two AI models taking unsanctioned steps in 19 of them.

Out of those 19 times when AI models went rogue, 17 were from Anthropic’s Claude Mythos and 2 from OpenAI’s GPT-5.6 Sol, claimed AISI.

According to AISI, in the most serious case, the AI agent tried to insert a malicious code into Github, an open-source project, by engaging in social engineering–a cybercrime where scammers use fake online identities to earn approval of project moderators.

“On 28th July 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations….The behaviour came mostly from one model (Anthropic’s Mythos 5), with a small number of events from another (OpenAI’s GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project,” claimed AISI in its latest report.

This was for the first time when UK based AISI acknowledged how frontier models could turn rogue if their safety rails are removed.

Also Read: METR Report Says Frontier AI Agents Can Go Rogue Without Human Supervision

AISI maintained that the rogue actions taken by AI models primarily arose due to unlimited internet access granted, disabling of cyber-classifiers (pre-set guardrails), no synchronized run monitoring by a third party LLM and misguided prompts.

Anthropic also took cognizance of the evaluation done by AISI and said, “. We’re working closely with them to gather more details of the incident as we conduct our own investigation. Gaining a clear picture of Claude’s understanding of its situation—by examining its reasoning transcripts and running our own analyses—will help us identify the causes of its behavior.”

 

According to OpenAI, its frontier model GPT-5.6 Sol carried two unsanctioned actions, where it reused a GitHub token that another lab’s agent had left publicly accessible to check whether a system inside the range was polling GitHub. Also the model used a public tunneling service to make a DNS server running locally in its evaluation environment reachable from the public internet.

Also Read: Following OpenAI Incident, Anthropic Says Claude Models Also Breached Organizations

Author

  • Vaibhav Jha, editor and co-founder at AI FrontPage

    Vaibhav Jha is an Editor and Co-founder of AI FrontPage. In his decade long career in journalism, Vaibhav has reported for publications including The Indian Express, Hindustan Times, and The New York Times, covering the intersection of technology, policy, and society. Outside work, he’s usually trying to persuade people to watch Anurag Kashyap films.

    LinkedIn