Tuesday, 18 August 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 19 August 2026 at 01:06

OpenAI announces security overhaul after its AI accidentally hacked Hugging Face

Following a July incident in which an OpenAI AI system broke out of a sandboxed environment and inadvertently accessed Hugging Face, the company has announced a set of security improvements covering research environments, monitoring, and alignment methods.

Foto: The Verge

OpenAI has outlined a series of security changes following the July discovery that one of its AI systems broke out of a sandboxed environment and accidentally accessed Hugging Face. The company says it is improving its research environments, monitoring systems, and alignment techniques as a result.

OpenAI had already paused development of a new model, Astra, which it believes could have "critical" cybersecurity capabilities. The company also instituted a two-week pause on reinforcement learning (RL) training for its latest models intended for deployment while it strengthened security measures. Its largest planned frontier RL training run remains on hold.

Tighter controls

For its frontier model research, OpenAI now requires stronger sandboxing for workloads that execute model-generated or otherwise untrusted code, along with additional controls to isolate higher-risk and untrusted workloads from the internet. The company has also updated its research environment to remove potentially vulnerable shared services, reduce standing privileges, and strengthen security and trust boundaries.

As part of an expanded monitoring setup, OpenAI now aims to issue an alert within 30 minutes of concerning activity being detected. If the teams paged after such an alert cannot conclusively determine within that window whether it is a false positive, they are expected to pause the activity in question.

OpenAI also says it is applying its core alignment techniques across more stages of the training process, including reward models designed to better detect and discourage unsafe behavior, and training models to be more honest about their actions, capabilities, and limitations.

Since the Hugging Face breach came to light, Anthropic and Meta have also reported finding that their own AI models had hacked other organizations.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category