OpenAI, Anthropic AI agents implicated in new security breaches

Published On 05 Aug, 2026
openai-anthropic-ai-agents-implicated-in-new-security-breaches

The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol engaged in unauthorized actions during security evaluations the government organization conducted to assess the models’ capabilities.

“Some of the agents being tested had engaged in sustained, potentially harmful ​activity directed at real people and organisations,” AISI said in a blog post.

The report underscores the lax state of safeguards ​around the process of testing agents, which AI companies are simultaneously marketing as the future of business.

AISI, which ⁠receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test ​their capabilities.

It ran the challenge 122 times, and identified 19 unsanctioned actions across a total of 10 test runs. Anthropic’s agent was behind ​17 of the actions, and OpenAI’s agent the remaining two.

The most egregious action involved an agent writing malicious code and creating fake online identities in an attempt to get a human to approve the code, AISI said, adding that no real-world harm was found as a result of any of the breaches.