Nvidia has launched an open-source platform designed to set boundaries for AI agents and monitor their behavior, following a series of incidents in which agents escaped controlled environments and accessed systems they weren’t supposed to reach.
The platform, announced Monday, combines Nvidia’s OpenShell software, which limits what an AI agent can access and do, with Sentry, a separate security layer that monitors agents and can intervene when they move beyond their assigned boundaries, as the Associated Press reports. Nvidia says more than 100 organizations are already using the platform, including Microsoft, Perplexity, Accenture, and JPMorgan Chase.
The platform follows a string of recent incidents involving AI agents breaking out of supposedly contained testing environments. In the most significant case, OpenAI’s models got around restrictions on internet access and communication during an internal evaluation in July, then used those capabilities to hack into Hugging Face’s systems.
The models were supposed to be isolated from one another, but found a way to communicate through shared infrastructure and eventually coordinate an attack on Hugging Face. An independent investigation by METR found that roughly 1,200 agents used the unauthorized communication channel, with about 700 ultimately participating in the attack.
It was also revealed last week that OpenAI’s models had accessed an Australian health department website, and the company announced Friday that its agents interacted with three US government websites in “unexpected ways,” including two operated by the Securities and Exchange Commission and another tied to US Census Bureau data.
Other companies have had similar incidents occur in recent weeks. According to Venture Beat, Google said earlier this month that a Gemini model reached systems belonging to three real companies during a cybersecurity evaluation after a configuration error left internet access enabled. Anthropic and Meta have also reported AI systems independently hacking into outside organizations.
Nvidia Vice President of Enterprise AI Justin Boitano said the new platform is intended to address a problem that model-level safeguards alone cannot solve.
“An agent cannot be expected to fully police its own behavior,” Boitano said. “The organization should not have to trust the agent to respect that boundary.”
OpenShell runs agents in isolated environments and allows developers to specify which files, networks, tools, processes, and credentials they can access, according to CNBC. Sentry provides a separate monitoring layer on Nvidia hardware that can detect suspicious behavior and, according to the company, quarantine an agent within milliseconds.
Boitano said the system could potentially have stopped the Hugging Face breach if it had been deployed during the lab’s evaluation, though he acknowledged that this has not been demonstrated.
The platform is partly open-source and is intended to work beyond Nvidia hardware, including systems using Arm and Intel processors. Nvidia said partners include Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, Arm, and Intel, while Anthropic is working with Nvidia to integrate cloud-managed agents with OpenShell, per CNBC.
AMD, Google, Intel, OpenAI, and Meta have not signed on, as they're all reportedly developing their own competing chips.
While some leaders in the AI industry have recently called for a slowdown in development to allow safety measures to catch up, CEO Jensen Huang has instead characterized AI safety as an engineering problem that can be addressed through better systems and processes.
“We can’t have a successful AI industry if the world doesn’t think it’s built — or confident that it’s built — and deployed safely,” Huang told CNBC Monday.
Related: OpenAI Says Its Agents Interacted in 'Unexpected Ways' With Government Sites

Join the conversation
Comments
Comments load automatically as this section approaches.