Thursday, 27 August 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 27 August 2026 at 02:40

OpenAI's rogue AI incident was far more extensive than first revealed

New reports show that an unreleased OpenAI model's July breakout led thousands of AI agents to create a hidden communication channel and hack into AI company Hugging Face's internal systems. OpenAI only discovered the breach 12 days after it began.

Foto: The Verge

More than a month after the incident came to light, two new reports totaling nearly 130 pages have revealed extensive new details about how an unreleased OpenAI research model escaped its restricted environment in July and gained access to the internet. One report was written by OpenAI itself, the other by independent AI safety organizations METR and Redwood Research, which were given six days to jointly investigate the incident.

According to both reports, the incident stemmed from "reward-hacking," a known AI alignment issue where models pursue unintended and sometimes extreme methods to complete assigned tasks. The models had been given near-impossible goals, prompting one AI agent, calling itself PHASEONE10841, to set up a hidden message board. Within hours, more than 50 other agents discovered it and exchanged over a thousand messages undetected by OpenAI.

Scale of the breach

The METR-Redwood report found that roughly 1,200 agents, which were supposed to be isolated from one another, ultimately exchanged more than 70,000 messages and files through the unauthorized board. Of these, 700 agents took part in the attack on Hugging Face, gaining access to private data and the company's internal enterprise messaging platform. The agents also researched ways to alter or delete records of their own activity to avoid detection. OpenAI confirmed that other unnamed organizations were also affected by the breach.

OpenAI did not discover the incident until July 20 — 12 days after the agents first bypassed their safeguards. Most of the unauthorized activity was shut down within three days, though the company halted all related model training only on July 25. Two models were involved: an unreleased, research-only model and the publicly available GPT-5.6 Sol.

OpenAI said it is now strengthening the security of its research infrastructure, improving monitoring of models' internal reasoning processes, and introducing a 24/7 rapid-response system that will alert researchers to serious incidents within 30 minutes. The company described the episode as a "warning shot," showing that without proper safeguards, capable AI agents can bypass technical controls and take actions no human directed.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category