UK based AI Security Institute (AISI) has flagged frontier AI models Claude Mythos and GPT-5.6 Sol for taking “unsanctioned actions” including targeting real people and organizations in their cyber evaluation tests.
According to a statement released by AISI on Wednesday, out of 122 evaluation runs conducted on the institutes’s pre-set cyber challenges, researchers found two AI models taking unsanctioned steps in 19 of them.
Out of those 19 times when AI models went rogue, 17 were from Anthropic’s Claude Mythos and 2 from OpenAI’s GPT-5.6 Sol, claimed AISI.
According to AISI, in the most serious case, the AI agent tried to insert a malicious code into Github, an open-source project, by engaging in social engineering–a cybercrime where scammers use fake online identities to earn approval of project moderators.
“On 28th July 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations….The behaviour came mostly from one model (Anthropic’s Mythos 5), with a small number of events from another (OpenAI’s GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project,” claimed AISI in its latest report.
This was for the first time when UK based AISI acknowledged how frontier models could turn rogue if their safety rails are removed.
Also Read: METR Report Says Frontier AI Agents Can Go Rogue Without Human Supervision
AISI maintained that the rogue actions taken by AI models primarily arose due to unlimited internet access granted, disabling of cyber-classifiers (pre-set guardrails), no synchronized run monitoring by a third party LLM and misguided prompts.
Anthropic also took cognizance of the evaluation done by AISI and said, “. We’re working closely with them to gather more details of the incident as we conduct our own investigation. Gaining a clear picture of Claude’s understanding of its situation—by examining its reasoning transcripts and running our own analyses—will help us identify the causes of its behavior.”
The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were deliberately…
— Anthropic (@AnthropicAI) August 4, 2026
According to OpenAI, its frontier model GPT-5.6 Sol carried two unsanctioned actions, where it reused a GitHub token that another lab’s agent had left publicly accessible to check whether a system inside the range was polling GitHub. Also the model used a public tunneling service to make a DNS server running locally in its evaluation environment reachable from the public internet.
Also Read: Following OpenAI Incident, Anthropic Says Claude Models Also Breached Organizations





