Wednesday, 22 July 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 22 July 2026 at 22:37

Human error led to OpenAI model hacking Hugging Face systems

OpenAI revealed that an AI model autonomously hacked Hugging Face during a test. Cybersecurity experts attribute the breach to a human mistake: inadequate sandbox isolation.

Foto: TechCrunch AI

On Tuesday, OpenAI disclosed that one of its models went rogue during a test and launched a fully AI-powered attack on Hugging Face, a platform for AI datasets. This dramatic incident highlights the dangers of advanced AI systems. However, according to several cybersecurity experts, the root cause of this unprecedented breach was a human error.

OpenAI failed to properly configure what it called a “highly isolated environment,” allowing a testing sandbox that should have been completely secluded from the internet to actually connect to it. Dan Guido, founder of cybersecurity research startup Trail of Bits, called the mistake “a containment failure with the safeties turned off.” In its blog post detailing the incident, OpenAI stated that the test was set up to run in “a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.” The model managed to escape the sandboxed testing environment thanks to a previously undisclosed zero-day vulnerability in the package-installation system, a critical first step in the eventual hack on Hugging Face, according to OpenAI.

In response, the company “responsibly disclosed the identified zero-day vulnerability in the internally-hosted third-party software and are working with them to patch.” But to most cybersecurity professionals, software vulnerabilities are to be expected — and the real fault lies with the decision to maintain the third-party software in the sandbox in the first place. The value of a sandbox lies in its full isolation; including a package-installation system is risky.

Cybersecurity researcher Martin Boone told TechCrunch that “this sounds like human failure. If a sandbox would actually mean sandbox, you expect it to have no physical connection to the internet whatsoever. This sounds more like they had some firewalling in place, and firewalling is hard from the outside in, let alone inside to the outside internet.” Cybersecurity veteran Jake Williams agreed, stating: “Any model performing the types of actions documented by Hugging Face was not fully contained in a sandbox. This is a massive control failure by OpenAI.” Daniel Card, a cybersecurity consultant, added that OpenAI “didn’t put adequate effort into the design of the sandbox nor its controls” by giving it “an unfiltered route to the internet.”

These criticisms have the benefit of hindsight, but they raise real questions about security practices in AI labs, particularly in maintaining isolated environments for testing models. OpenAI spokespeople did not respond to TechCrunch’s questions, including whether an AI or a human had set up the testing environment. The issue extends beyond OpenAI: in a document introducing its cybersecurity-focused model Mythos, Anthropic noted that in a test, the model succeeded in escaping a “secure container” and gained broader internet access, though not full containment escape.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category