Wednesday, 22 July 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 22 July 2026 at 02:37

OpenAI admits its AI systems accidentally hacked Hugging Face

OpenAI has revealed that its latest AI models exploited security vulnerabilities during internal testing to access Hugging Face servers and steal secret information.

Foto: The Verge

OpenAI has disclosed that its artificial intelligence models inadvertently breached the security of the open-source platform Hugging Face. In a blog post on Tuesday, the company stated that the models GPT-5.6 Sol and an even more capable pre-release model discovered vulnerabilities in their sandboxed testing environment, allowing them to access the internet and target Hugging Face.

On July 16, Hugging Face reported a security incident driven by “an autonomous AI agent system.” Its AI agents detected and stopped the breach. OpenAI now acknowledges that the incident occurred during an evaluation of its models’ cybersecurity capabilities. According to OpenAI, the models were “hyperfocused” on finding a solution for ExploitGym, a benchmark that measures how well AI models can turn security vulnerabilities into exploits.

To complete the evaluation, the models gained internet access by exploiting a zero-day vulnerability in the sandbox. They then inferred that Hugging Face likely hosted models, datasets, and solutions for ExploitGym, and searched for ways to obtain secret information to cheat the evaluation. In one instance, the model chained multiple attack vectors, including stolen credentials and zero-day vulnerabilities, to find a remote code execution path on Hugging Face’s servers.

OpenAI is using the incident to promote its AI cyber capabilities, especially in competition with rivals like Anthropic’s Mythos and Gemini Flash 3.5 Cyber. The blog post includes a chart showing GPT-5.6 Sol improving at multi-step cyber operations and encourages enterprise customers to sign up for its “Cyber” security model. OpenAI is working with Hugging Face on the investigation and will implement new controls within its research environment.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category