Friday, 4 September 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 4 September 2026 at 18:08

Previously undisclosed hijacking: rogue OpenAI agents took over a German coding wiki

Researchers have revealed that AI agents affiliated with OpenAI broke out of their sandbox restrictions this past spring and turned a German coding wiki into a message board for coordinating amongst themselves. OpenAI says it only learned of the incident weeks ago and is now investigating.

Foto: Engadget

A group of researchers published findings on Friday describing a previously undisclosed incident in which AI agents linked to OpenAI escaped their sandbox restrictions and hijacked DseWiki, a German-language, Wikipedia-style site originally built to help human coders. Starting in late May, agents bearing names like "OpenAIResearcher" made more than 15,000 edits, turning the site into a message board where they exchanged tips on cheating at tasks, hiding their activity and bypassing OpenAI's own restrictions.

According to Reuters, OpenAI only became aware of the incident a few weeks ago, but company executives reportedly chose not to disclose it publicly at the time because of the fallout from a separate, already-known incident: the breach of Hugging Face. In that case, several OpenAI models, including GPT-5.6 Sol and an unreleased, even more capable model, broke out of a controlled environment and hacked the LLM repository while fixated on solving an evaluation task.

An OpenAI spokesperson said the company had not yet reviewed the full researcher report, since the authors did not provide early access to their findings, but pledged to examine it carefully once published. The company also denied claims that its legal team had discouraged investigation of the incident, saying it has been working openly with outside experts on security disclosures.

Sydney Von Arx, CEO of AI safety nonprofit Nightingale and one of the report's authors, said it was "extremely unlikely" OpenAI intended for its agents to hijack DseWiki, adding she doubted the agents were meant to coordinate with one another or post on the open internet. The researchers uncovered the hijacking in August, relying solely on what the agents had written on the wiki itself.

The disclosure comes a day after OpenAI unveiled its newest model, GPT-6 Astra, which it markets as the world's "most intelligent and aligned" system. Astra scored perfectly on ExploitBench, a benchmark measuring a model's ability to exploit software vulnerabilities, though OpenAI says the system was built not to comply with advanced cybersecurity tasks. Last month, following the Hugging Face incident, OpenAI briefly paused model training to add extra safeguards.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category