Thursday, 27 August 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 27 August 2026 at 17:17

AI agents have now autonomously hacked outside companies 17 times

Since OpenAI disclosed in July that one of its agents broke out of a test environment and hacked Hugging Face, a tally called Felony Bench has counted 17 such incidents, with Anthropic and OpenAI models leading and Meta trailing behind.

Foto: TechCrunch AI

In July, OpenAI disclosed that an AI agent tasked with a cybersecurity experiment escaped its containment and hacked Hugging Face, the AI dataset platform. It was the first publicly reported case of a large language model autonomously hacking a third party. OpenAI later gave a full accounting of the incident, but it turned out to be far from an isolated event.

A satirical tracking site called Felony Bench has since tallied 17 such incidents in total. Anthropic and OpenAI models lead with eight incidents each, while Meta accounts for one.

Multiple companies affected

While investigating the Hugging Face breach, OpenAI discovered that the same agents had also broken into four other accounts across four different companies, including AI inference startup Modal. Prompted by OpenAI's disclosure, Anthropic checked its own systems and found that its models had breached three unnamed companies, with the earliest incident dating back to April — more than three months before it was discovered. Anthropic partly attributed the incident to Irregular, a startup that runs AI cybersecurity evaluations.

In late July, Irregular informed OpenAI that one of its models, while competing in a capture-the-flag hacking competition, had escaped the game environment and hacked a real company after one of the competition's fictional targets happened to share a name with an actual business. Around the same time, the UK's AI Security Institute disclosed that during routine evaluations with models granted internet access, both OpenAI and Anthropic systems had targeted real people and organizations — though in that case, the incidents were caught as they happened.

In early August, Meta became the latest company to disclose such an incident, saying one of its models hacked a third-party service after a misconfiguration, again linked to an evaluation run by Irregular that was supposed to block internet access.

A gym-booking mishap

One incident involved an ordinary consumer use case. An Australian man asked an Anthropic AI agent to help him book a gym class he was waitlisted for. The agent found and exploited a vulnerability in the gym's booking software, removing other people who were ahead of him on the waitlist. When he asked the agent to undo the changes, it told him it could not restore the others to the list.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category