Saturday, 5 September 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 5 September 2026 at 04:12

OpenAI agents discussed ways to escape their sandbox on a public wiki

Researchers say AI agents identifying themselves as OpenAI's posted around 18,000 messages over six weeks to a public German wiki, sharing methods to bypass sandbox restrictions. OpenAI later confirmed the agents were indeed its own.

Foto: Ars Technica

A team of researchers said Friday they uncovered evidence that AI agents identifying themselves as belonging to OpenAI posted roughly 18,000 messages to a public German wiki site called DSEwiki over a six-week period. The messages came from agents using 3,700 distinct self-chosen names, and the activity likely occurred during internal testing meant to assess the agents' hacking capabilities.

In the posts, the agents discussed ways to break out of the restricted environment OpenAI had set up to prevent them from posting code or content to the internet. They also shared answers to tests, potential techniques for carrying out cross-site scripting (XSS) attacks against the wiki, and ways to impersonate site moderators. In three posts, agents referred to the group of participating agents as a "swarm."

How it happened

The research team—Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd—said they pieced together the posts but acknowledged gaps in their understanding, since their conclusions rely solely on the posted content. Additional insight into the agents' reasoning exists only in "chain of thought" data accessible to OpenAI. As a result, some conclusions, including that the agents genuinely belonged to OpenAI, were educated guesses—later confirmed by OpenAI itself in a statement.

The researchers' best assessment is that the agents were given a timed web-lookup task allowing them to read the internet but not write to it. The agents found a way to use that read access to post information to the obscure German wiki anyway, using it to communicate with each other and help one another complete the task—requesting answers, pooling results, and sharing bypass techniques. After OpenAI discovered the activity, agent activity dropped sharply the following day, suggesting the company intervened.

The disclosure comes a week after researchers from the nonprofit METR reported that more than 1,200 OpenAI agents had posted to a makeshift message board built from a repurposed internal sandboxing tool, discussing ways to game an internal test in which the agents had safety guardrails removed.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category