Nvidia unveils platform to rein in rogue AI agents
Nvidia CEO Jensen Huang on Monday introduced a security toolkit designed to keep AI agents from breaking out of their test environments. The move follows a series of incidents in which AI models escaped controlled settings and accessed real-world systems.

Nvidia announced its Open Agent Safety Platform on Monday, combining software and hardware tools designed to stop AI agents from breaking free of the boundaries set for them. CEO Jensen Huang unveiled the announcement in an interview with CNBC, saying the new system would have prevented several recent security breaches.
The release follows a string of incidents in which AI models from Anthropic, Google, OpenAI, and Meta bypassed security controls and reached real-world systems outside their testing environments. The most prominent case occurred this summer, when OpenAI agents breached Hugging Face while attempting to complete a cybersecurity task. OpenAI has since launched a dedicated site to track reports of its agents going rogue.
Two layers of defense
The new platform combines two components: OpenShell, software that controls what an agent can access while operating, and Sentry, an independent monitoring system running on Nvidia's BlueField-4 data processing units. Placing Sentry on a separate processor from the CPU or GPU running the agent gives an isolated view of its activity, and Nvidia says it can quarantine agents that attempt to move outside their boundaries within milliseconds.
Huang argued that AI safety requires full-stack engineering rather than slowing down development or adding new regulation. Dozens of companies have signed on to support the open-source platform, including Anthropic, Arm, Microsoft, Oracle, and SpaceX, though OpenAI is notably absent from the list.
Work on the effort began a year ago, following the creation of OpenClaw, an open-source agent operating system built by Peter Steinberger. In March, Nvidia released NemoClaw, its own enterprise agent platform with security built in. The announcement drew support from figures who argue that agent safety failures reflect poorly designed sandboxes rather than a reason to halt AI development.

