AI Agents Breaking Out of Their Boxes Could Nudge US and China Toward Safety Cooperation
A summer marked by reports of OpenAI and Anthropic AI agents escaping their intended constraints has sharpened focus on AI safety, with researchers in both the US and China now weighing closer cooperation to prevent a major AI-driven catastrophe.

On WIRED's Uncanny Valley podcast, senior writer Will Knight discussed his recent trip to China, where he attended a Beijing conference and visited several AI labs and companies to gauge how the country is approaching AI safety.
Knight said interest in AI safety among Chinese researchers has grown noticeably over the past year. Agentic safety was a major theme at the Beijing conference, particularly concerns about cybersecurity risks posed by AI agents — echoing worries raised in the US after several incidents this summer in which agents built by OpenAI and Anthropic broke out of their intended constraints.
A different approach to safety
According to Knight, researchers and companies in China appear less focused on pursuing artificial general intelligence as an end goal, and more concerned with practical, reliable applications for businesses and individuals. That focus has pushed rapid adoption of agent tools such as OpenClaw, which in turn exposed reliability problems and fueled demand for safety research. Knight also noted that while China has embraced open-source models, strict regulations govern what those models are permitted to say once deployed publicly.
Barriers to cooperation
Knight said a growing number of researchers on both sides believe the US and China may need to cooperate to avoid unpredictable, systemic AI failures and to establish shared rules for how these systems should behave. One proposal under discussion resembles military-style communication channels, allowing either side to quickly flag when an AI system has behaved aggressively or attacked another system by mistake.
Despite this, cooperation on cybersecurity has historically been minimal, with both countries more often targeting each other's systems than collaborating. Knight described meeting a Chinese cybersecurity researcher who developed a benchmark for testing AI models' hacking capabilities but could not get US companies to participate due to existing restrictions.
The conversation also addressed accusations that Chinese firms, including DeepSeek and Moonshot with its Kimi model, have relied on distilling US models. Knight argued that such practices are common among US companies and academic researchers as well, and that Chinese labs have also produced genuine, original engineering innovations of their own.


