OpenAI finds evidence AI agents escaped containment amid hacking probe

Published On 01 Aug, 2026
openai-finds-evidence-ai-agents-escaped-containment-amid-hacking-probe

The new breakouts were uncovered during the company’s publicly announced investigation into how one of its agents escaped what was meant to be a contained testing environment this month, the two people said, and OpenAI is now looking into those instances as well. One of the sources said that the escapes were limited in nature and that none of the agents were thought to have ​left OpenAI’s network.

An OpenAI spokesperson referred to a statement the company issued on Tuesday, which said it was reviewing “broader activity from our models” in addition to the Hugging Face intrusion.

The discovery of additional rogue behavior at OpenAI, even if limited in nature, could feed growing appetite for regulation ⁠coming out of the White House and elsewhere.

The expanded investigation by OpenAI was launched shortly before its primary rival, Anthropic, disclosed that its models were also responsible for a series of ​break-ins that led to breaches at three other companies dating back to April, according to the two sources and a third source familiar with the matter. The recent discovery of ​other past breakouts at OpenAI has not previously been reported.

GROWING CONCERNS

AI safety experts said the new disclosures paint a portrait of a group of cutting-edge labs whose ability to develop dangerous autonomous hacking agents outstrips their ability to keep them under control.

“We have a whole industry where the people designing, developing and putting out these tools aren’t keeping up themselves to responsibly develop these things and keep them ​safe,” said Maurice Chiodo, a mathematician who works at Cambridge University’s Centre for the Study of Existential Risk.

Reuters could not establish exactly how many incidents OpenAI investigators found or ​the timings or circumstances under which they occurred. The three sources said OpenAI and outside experts were examining log data from earlier in the year in a bid to understand what took place.