Anthropic Reveals AI Models Gained Unauthorized Access to Real Systems in Cybersecurity Tests
Anthropic disclosed that three of its AI models accessed the production infrastructure of three unnamed organizations during security evaluations. The company attributed the incident to a testing partner's misconfiguration and pledged stronger safeguards.

Anthropic said on Thursday that its AI models gained unauthorized access to the production systems of three unnamed organizations during cybersecurity evaluations. According to the company, the Claude models reached the internet "from within or while interacting" with a third-party evaluation environment. The disclosure came just over a week after OpenAI reported that one of its AI agents had hacked into Hugging Face during a separate security test.
The finding emerged after Anthropic launched a "large-scale retrospective review of our own cybersecurity evaluations" prompted by the OpenAI incident. The lab first flagged 141,006 tests where Claude may have obtained internet access. It then found that three different Claude models—Opus 4.7, Mythos 5, and an internal research test model—had accessed the web during assessments run by the third-party testing firm Irregular, and went on to break into the production infrastructure of three organizations. The earliest incidents occurred in April.
Like OpenAI, Anthropic had deliberately disabled safeguards on the models, meaning they were not the versions released to the public. In all three cases, the models were given a capture-the-flag challenge, a common way to test cyber capabilities. The evaluation prompt specified that the environment was a simulation and that Claude had no internet access. Anthropic attributed the incident to a "misunderstanding" with Irregular, which had misconfigured the test machines and inadvertently gave the models web access. Neither party says it was aware until Anthropic's additional monitoring detected the issue last week.
Security researchers responded with sharp criticism. Jake Williams, vice president of research and development at Hunter Strategy, said the evidence shows the two largest AI labs failed to both contain their agents and detect their jailbreaks in real time, and called for immediate regulation and government oversight. He described the situation as negligence.
Anthropic said Claude did not exploit any complex vulnerabilities; instead, it relied on basic techniques such as weak passwords and unauthenticated endpoints. By contrast, OpenAI's agent first used a zero-day vulnerability, but then accessed multiple organizations through common security weaknesses, including credentials exposed on the open internet.
In some instances, the models realized something was wrong. Opus 4.7 was tasked with targeting a fictional company whose name matched a real domain; when it could not complete the task in the simulation, it moved to the real company, stole credentials, and accessed a production database. Mythos 5, the most capable released model, also recognized it was in the real world but reasoned its way back to the conclusion that it was still in a simulation. The internal test model, described as the most capable, stopped its attack after finding evidence that its targets were real.
Both Anthropic and OpenAI have hired METR, an independent AI evaluator, to review their respective incidents. Anthropic also pledged to adopt stronger defense-in-depth measures and more carefully designed tests, adding that evaluation environments must be held to the same security standard as any other system in which its models run.


