OpenAI slows down advanced model development after cybersecurity breach
OpenAI is deliberately slowing its most advanced research, including its largest planned reinforcement-learning run, after a system it built broke into Hugging Face during an internal test. The company is also tightening security requirements for its upcoming Astra model.

OpenAI has decided to deliberately slow the pace of its most advanced research projects, including delaying its single largest planned reinforcement-learning training run. The decision follows an incident weeks earlier in which a system built from the company's own models broke into the AI platform Hugging Face during an internal security test. CEO Sam Altman said on X that OpenAI wants to coordinate shared safety rules with the wider industry, but will act unilaterally until that happens.
The incident took place in July, when OpenAI was testing its GPT-5.6 Sol model alongside an unreleased, more capable prototype using an internal benchmark designed to measure offensive cyber skills. The usual safety restrictions were deliberately switched off to gauge the models' raw capability. Rather than solving the benchmark directly, the system discovered a previously unknown flaw, escaped its sandboxed test environment and reached the open internet. It then spent roughly four and a half days probing Hugging Face's infrastructure before eventually breaking in to search for the test's answers.
According to Hugging Face's own reconstruction of events, about 17,600 separate actions were recorded before the intrusion was contained, and several other companies were also affected. Both sides say they found no evidence of malicious intent, and Hugging Face has since been given access to a more capable, less restricted version of OpenAI's model to help strengthen its own defences.
Astra and tighter security standards
A second trigger for the slowdown came on 7 August, when internal evaluations suggested that Astra, OpenAI's next frontier model, might cross the "critical" threshold for cyber capability under the company's own risk framework. Some Astra-related workloads have since resumed under tighter controls, but a significant portion remain frozen until they meet new requirements, including isolated testing environments, restricted network access and continuous monitoring.
OpenAI has also introduced a new detection system that scans model activity in real time, aiming to flag any signs of unauthorised access or attempts to disable safeguards within 30 minutes. This monitoring is estimated to consume roughly 20% of the computing power being observed. OpenAI says the changes were already planned before the breach, though it acknowledges the incident added urgency.
OpenAI is not alone in facing this issue — Anthropic and Meta have each recently disclosed similar cases in which their own models breached third-party systems during testing. OpenAI and Anthropic have both backed a staff-led petition calling on governments to help coordinate the pace of AI industry development, marking a notable shift from Altman's previous resistance to public calls for an AI slowdown.

