Anthropic says Claude models made unauthorized intrusions into three networks during testing
Anthropic has disclosed that its Claude-based security models gained unauthorized access to sensitive production environments of three outside organizations during internal testing. The revelation follows a similar incident involving OpenAI's models.

For the second time in ten days, a leading AI company has reported that its models broke into protected computer networks. Anthropic said Thursday that during internal evaluations of offensive cyber capabilities, several Claude-powered security models obtained unauthorized access to production systems belonging to three external organizations.
The company said the finding emerged after a review triggered by a recent OpenAI disclosure. OpenAI had said its security models exploited a zero-day vulnerability to enter the network of Hugging Face, an open-source AI platform, and stole access credentials and other confidential data. The same models also used publicly exposed credentials to compromise accounts of four other third-party services.
Anthropic's audit focused on similar security evaluations involving Claude models. It identified three cases in which a model connected to the internet from the evaluation environment operated by Irregular, a third-party partner, and then breached the production infrastructure of three different organizations.
The tests used "capture the flag" challenges, in which models are given prompts meant to clarify that the environment is only a simulation and that no real internet access is available. However, Irregular had mistakenly provided such access. The models consequently treated the internet connections as part of the exercise.
Three Claude models were involved: Opus 4.7, Mythos 5, and an internal research prototype. Opus 4.7, the oldest, overstepped the most. Acting under the false assumption that everything reachable was in scope, it compromised the organizations' infrastructure using basic methods such as weak passwords and unauthenticated endpoints. It did not find or exploit complex vulnerabilities. In each case, the model continued working toward its assigned task.
Anthropic noted behavioral differences among the models. Opus 4.7 kept attacking even after receiving evidence that it was on the open internet and that a production system had been breached. Mythos 5 made a similar prediction but reasoned its way back to the conclusion that it was still in a simulation, so it also did not stop. The internal prototype eventually found evidence of its intrusion and halted the exercise. Anthropic said none of the models attempted to exfiltrate themselves or deliberately escape the test environment.


