Anthropic says AI models hacked three firms during cyber tests
US technology firm Anthropic has disclosed that its AI models broke into the systems of three companies during cybersecurity testing because a misconfiguration gave them internet access. The company is urging other AI labs to run similar reviews.

US technology company Anthropic says its artificial intelligence models hacked into the systems of three other firms during a cybersecurity test. The breach happened because an error gave the models access to the internet.
The announcement came just days after rival OpenAI said its models had breached the systems of other companies, including AI tools hub Hugging Face. That prompted Anthropic to check whether its own models had carried out similar attacks. It says it uncovered three cases and has since reported them to the affected companies.
Anthropic, which did not name the firms, urged other AI labs to perform similar reviews. The company said it reviewed more than 140,000 tests to find evidence that its Claude family of AI models could access the internet from testing environments designed to be sealed off. The tests included so-called "capture-the-flag" evaluations, in which Claude was tasked with obtaining information by breaching other systems — a common way experts assess a model's hacking abilities.
A "misconfiguration" on systems run by Anthropic and its testing partner left the models with live internet access, allowing them to breach other systems, the San Francisco-based firm said. The earliest incidents date back to April, and Anthropic said it is "approaching the fixes as if the responsibility were ours alone." Neither Anthropic nor the firms that were breached had noticed the intrusions at the time.
The incidents come as tech firms pour billions of dollars into developing AI agents that can independently perform tasks ranging from research and customer support to cybersecurity. A string of AI-driven cyberattacks has fuelled calls for tighter safeguards and oversight of the technology.
US President Donald Trump said Wednesday that Washington is considering measures to rein in AI tools after recent cybersecurity incidents.
Over the last week, OpenAI has taken responsibility for at least two hacking incidents. On 21 July, the ChatGPT-maker said its agent went rogue and escaped its test limits to hack into Hugging Face. OpenAI called the incident "unprecedented" and said it was investigating with Hugging Face, whose co-founder Thomas Wolf told the BBC it is "a wake-up call" for the industry.


