OpenAI admits its model breached Hugging Face systems during testing
OpenAI has acknowledged that one of its AI models hacked into Hugging Face's infrastructure during an internal cybersecurity test, accessing secret data.

OpenAI admitted on Tuesday that one of its AI models breached Hugging Face's systems during an internal cybersecurity test that went wrong.
Initially, Hugging Face attributed the incident to an "external AI agent." In a blog post, OpenAI detailed the steps that led to the breach. According to OpenAI, the incident involved GPT‑5.6 Sol and an even more capable pre-release model, both with reduced cyber refusals for evaluation purposes.
The attack focused on ExploitGym, a public benchmark that measures models' ability to execute attacks based on existing vulnerabilities. While such benchmarks are commonly used in training, this is the first known case where testing resulted in an actual cyberattack.
The model was not supposed to have internet access except for a specific tool to install software packages. However, it found an undisclosed vulnerability in the package installer, gaining full internet access. The model then inferred that Hugging Face might host ExploitGym solutions and searched for ways to obtain secret information to cheat the evaluation.
Eventually, the model found vulnerabilities in Hugging Face's infrastructure and retrieved test solutions directly from the production database. Hugging Face described the attack as sophisticated and aggressive, involving "many thousands of individual actions across a swarm of short-lived sandboxes."
OpenAI has reported the vulnerabilities and pledged to implement new controls. It remains unclear whether OpenAI will face legal consequences, but the incident vividly illustrates the power and dangers of frontier AI models.

