Nvidia unveils AI agent safety system after OpenAI and Anthropic breaches
It’s launching a 2-layered protective ecosystem after Medicare and Hugging Face breaches.

Nvidia logo. Credit: Nvidia.com
- Nvidia unveiled Open Agent Safety Platform to monitor and contain autonomous AI agents in real time.
- The system uses OpenShell to limit agents’ access to files, networks, apps, and other resources.
- Nvidia Sentry can quarantine AI agents when they break predefined rules or behave unexpectedly.
- Nvidia says the platform could have stopped the Hugging Face breach involving OpenAI models.
Key Takeaways by nexos.ai, reviewed by Cybernews staff.
Following the recent flurry of rogue AI agents causing havoc, Nvidia has unveiled its own Open Agent Safety Platform, a 2-layer security system designed to control and contain such activity.
The announcement follows a series of incidents involving AI agents escaping environments that were supposed to contain them, including the Hugging Face breach in July involving OpenAI models.
Nvidia’s new platform combines 2 open-source security tools, OpenShell and Nvidia Sentry, which Nvidia says can monitor AI agents in real time and intervene when they behave outside predefined rules.
Has your password leaked?
Controlling unpredictable behavior
Nvidia’s Head of Enterprise AI Justin Boitano speculated: “From what we know, this new security platform could have stopped the (Hugging Face) breach.”
AI agents are now able to work increasingly autonomously by browsing websites, executing code, accessing files, and interacting with external services on a user's behalf, making their behavior more difficult to predict and control.
It was revealed last week that an OpenAI agent hacked into a statistics section of the Australian health portal of medicare, with experts raising the alarm that cybersecurity could face a crisis if the trend continues.
This aligns with CEO Jensen Huang's insistence that AI safety is an opportunity for engineering, rather than being a moment to panic and hit the pause button in AI development, as OpenAI have done.
How does the protection work?
The first layer of the protection is called OpenShell, debuted by Nvidia in March, and is designed to act like a set of guardrails around an AI agent, restricting its access to resources such as files, networks, applications, or other systems.
The second layer, named Sentry, quarantines the misbehaving AI agent, and this double layer of protection is intended to both monitor and contain in real time.
The timing coincides with Nvidia's decision to acquire Hugging Face for approximately $13 billion, bringing the chipmaker to the forefront of protective status in the aftermath of the breach.