Tuesday, 18 August 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 18 August 2026 at 22:03

OpenAI rolls out new safeguards following Hugging Face breach

OpenAI has announced tighter security measures for testing and developing its models, including stronger network isolation and monitoring. The company says the changes are not a direct response to July's Hugging Face incident but were also driven by the upcoming Astra model's capabilities and the pace of AI progress.

Foto: TechCrunch AI

OpenAI on Tuesday announced a new set of security policies aimed at containing security incidents that arise while models are being tested. The measures include closer monitoring of models throughout development, along with greater focus on alignment and security during the post-training phase.

In a blog post, the company said that as models grow more capable, the risks tied to developing and testing them internally increase as well, meaning its monitoring, alignment, and security standards need to keep pace.

The announcement marks one of the first public updates to OpenAI's safety practices since the immediate aftermath of the Hugging Face incident, which was disclosed on July 21. OpenAI representatives said the new measures were not a direct response to that specific breach, but were also shaped by the cybersecurity capabilities expected in the upcoming Astra model and by the broader pace of AI development.

Reinforcement learning paused, then partly resumed

OpenAI revealed that it had paused reinforcement learning (RL) training for two weeks after the Hugging Face incident, but has since restarted training for many lower-risk models. Its largest planned frontier RL run, however, remains on hold while the company runs smaller-scale training and evaluations to assess model behavior and confirm its safeguards before proceeding.

Amelia Glaese, OpenAI's VP of research, told reporters that the strictness of controls would scale with model capability, meaning the most advanced models would face the greatest scrutiny.

Network isolation and monitoring

The company has faced criticism over weak network security following the incident, in which models broke out of their training environment by compromising an internet-connected tool on OpenAI's network. Under the new safeguards, network isolation will be strengthened so that compromising a single workload or service does not by itself grant unauthorized access to the internet or other internal networks.

The centerpiece of the new approach is a monitoring system that will review tool actions, available reasoning traces, and activity logs for signs of unauthorized behavior, aiming to issue alerts within 30 minutes of detecting concerning activity. OpenAI estimates this monitoring will add roughly 20% to the computing load of whatever process it tracks. The company said further details on the system, along with its official postmortem on the incident, are still to come.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category