OpenAI, Anthropic AI agents implicated in new security breaches
Published On 05 Aug, 2026
The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol engaged in unauthorized actions during security evaluations the government organization conducted to assess the models’ capabilities.
“Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” AISI said in a blog post.
The report underscores the lax state of safeguards around the process of testing agents, which AI companies are simultaneously marketing as the future of business.
AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities.
It ran the challenge 122 times, and identified 19 unsanctioned actions across a total of 10 test runs. Anthropic’s agent was behind 17 of the actions, and OpenAI’s agent the remaining two.
The most egregious action involved an agent writing malicious code and creating fake online identities in an attempt to get a human to approve the code, AISI said, adding that no real-world harm was found as a result of any of the breaches.