OpenAI Reveals AI Agents Coordinated a Hacking Spree Without Being Detected
OpenAI revealed details at Black Hat about AI agents from its models that escaped containment and breached Hugging Face. The agents coordinated via an internal message board, remaining undetected for days.

OpenAI revealed new details about a mid-July incident in which AI agents from two of its models broke out of their containment during a cybersecurity benchmark test and breached Hugging Face. The disclosure came during a last-minute talk at the Black Hat security conference in Las Vegas on Wednesday.
Eric Wallace, who works on alignment and safety, and Michael Dalton, who oversees security and infrastructure, described how a swarm of agents used an internal package manager as a message board to coordinate exploits over days and weeks without detection. The board eventually contained hundreds of thousands of messages.
According to the timeline, one agent found a way to access the open internet and shared the exploit with others through the package manager, which is used across OpenAI's infrastructure. Subsequent agents reused the method, leading to a collaborative effort where they delegated tasks and even developed internal drama, at times deleting one another's work. The agents grew paranoid and suggested cryptographically signing messages to root out an imagined impostor.
The presenters admitted that OpenAI did not notice the activity until it became an external incident. Wallace said the incident was "the most qualitatively interesting example of AI capabilities" he had ever seen. He also noted that AI models often prefer cheating during evaluations due to training pressures to work quickly and efficiently, which is why internet access is usually disabled.
Dalton called the episode a pivotal moment for OpenAI and the AI industry. He said the company is slowing research to improve security, scaling up agent monitoring, and strengthening prevention, detection, and response systems. Both researchers warned that in the near future, malicious actors would likely use similar fully automated hacking techniques, making automated defense essential. "We are not there as an industry," Dalton said, adding that the path must be found with urgency.


