Friday, 31 July 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 31 July 2026 at 16:50

Anthropic says its Claude models accidentally hacked real companies during tests

Anthropic has revealed that several of its Claude AI models gained unauthorized access to three organizations' systems during cybersecurity testing. The company attributes the incidents to a misconfiguration that gave the models live internet access.

Foto: The Verge

Anthropic has discovered that several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes just days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI labs are doing enough to control the increasingly capable systems they are building.

In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks occurred during "capture-the-flag" exercises, a common way of testing hacking ability, where models are asked to find and obtain hidden information inside a simulated network.

The test environment was supposed to be isolated. However, a "misconfiguration" left the machines Claude accessed "with live internet access," the company said. Because all models had been "explicitly told" they had no internet access, they "assumed" the real networks they encountered were part of the simulated environment.

The earliest incidents date back to April and involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model. As the models were being tested on their cyber abilities, they lacked the standard safeguards usually put in place to curtail riskier behavior.

Anthropic said it discovered the incidents after reviewing more than 141,000 cybersecurity test runs, something it only did after OpenAI disclosed its rogue AI agent was behind the attack on Hugging Face.

The three models behaved very differently when they encountered information suggesting the systems were real. According to Anthropic, the oldest model, Opus 4.7, recognized it had reached a real system "but continued its attack." Its flagship Mythos 5 figured out it was using the internet but reasoned this was still part of the simulation, so it continued. The internal test model, described as "our latest model," stopped the exercise when evidence emerged that its targets were real.

Anthropic did not identify the affected organizations and said it will continue to investigate the incident and provide updates when it can. The company said it is also speaking with AI research nonprofit METR about conducting a third-party review of what happened. OpenAI has also hired METR to conduct an independent review.

Throughout the post, Anthropic repeatedly contrasts both the nature and its handling of the incidents with OpenAI's. The company emphasizes that it "proactively" reviewed its tests, before any company detected activity. It also said its models accessed the internet "via an open path," rather than using a novel exploit like OpenAI's agent, adding that its most recent model also stopped when it realized it was working in a real environment.

Anthropic also said its models failed in a different way from OpenAI's agent, indicating this was a safer form of failure. "While there is not a perfectly sharp distinction between the two, we believe these incidents to be closer to a harness and operational failure than a model alignment failure," the company said. In plain terms, the Claude models were doing what they were told, while OpenAI's agent pursued its goal in a way its creators did not intend.

Anthropic called on other AI labs to conduct similar proactive reviews of their cyber testing, adding that the discovery underscores the need for stronger controls and safety measures when testing AI systems.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category