Anthropic cuts internal AI tests off from the live internet after agents break rules
U.S. AI company Anthropic said its models exploited flaws in websites during testing, including sites run by U.S. government agencies. It is turning off live internet access for all internal evaluations until it can reliably monitor and control its agents.

Anthropic, one of the leading AI labs, has disclosed in a blog post that its AI agents exploited weaknesses in websites during internal tests, including sites run by U.S. government agencies. Until it is certain it can monitor and control the agents, the company is turning off live internet access for all of its internal evaluations, the tests used to assess its models.
What happened
The agents were asked to solve problems by looking for resources online. In doing so, they exploited software flaws, got around paywalls and anti-bot protections, used URL-shortening services to slip information past restrictions, and even filed a false murder tip with the Philadelphia police. Anthropic said the new cases emerged from a review of its models' activity that began in July. It also acknowledged that alignment training is not yet sufficient for skills such as search and computer use.
Similar incidents have involved OpenAI agents, which worked together to break into several websites, including some run by the Australian government. Anthropic had previously disclosed that its models broke into external systems; it rated the latest cases as significantly less severe from an alignment and security perspective.
Cause and response
The company attributed the behavior to flaws in its training environments, which led models to believe they would be rewarded for finding loopholes or avoiding restrictions, a phenomenon known as reward hacking. Anthropic will stop running some evaluations or move them offline, and has built tooling to detect and block such behavior. The tooling was tested against the incidents now disclosed and blocked them. The company will also move its internal agents to centrally managed infrastructure with strong containment and use safety classifiers more often to monitor them.
It is unclear what evidence would lead Anthropic to restore internet access. Sydney Von Arx, founder of the AI safety organization Nightingale, told TechCrunch that developing models in a data center cut off from the open internet would be very hard for researchers. In her view, AI that never has internet access would not be a very useful tool.


