kras99 - stock.adobe.com
Nvidia launches AI safety platform after agent security breaches
The AI vendor's new system is designed to contain rogue AI agents as companies race to secure increasingly autonomous systems. But it has not been proven yet.
Nvidia on Monday launched an AI safety platform that the AI giant said can contain rogue agents within “milliseconds,” as concerns over unchecked AI systems mount.
The AI chip and model maker’s Open Agent Safety Platform enables developers to set their own safeguards for agents and prevent them from breaking out of permitted environments.
The development follows a spate of incidents involving AI agents from major tech vendors, including OpenAI, Anthropic, Meta and Google, that bypassed security controls to hack external systems.
The landscape has sparked new scrutiny of AI systems’ guardrails, with Anthropic CEO Dario Amodei calling at the beginning of this month for developers to slow the pace of advancement to keep models in check.
Nvidia said the incidents shared a common pattern, with agents circumventing safeguards built into the application layer to complete their assigned tasks. Its platform is designed to combat this problem, providing an additional security boundary at the infrastructure level.
While Nvidia CEO Jensen Huang has previously argued against those pushing to slow down AI and rejected the risk of human extinction from AI, the release of the agent platform doesn’t contradict his argument but rather furthers it, said Kashyap Kompella, CEO and founder of RPA2AI Research.
“Jensen has resisted the idea that making AI safer requires slowing its development,” Kompella said. “Nvidia’s approach is to keep advancing AI while building stronger controls around it.”
He said that Nvidia stands to benefit from this approach, so the new platform becomes both a response to the technical problem of AI security risk and a new commercial opportunity for Nvidia.
But Nvidia must ensure its platform integrates with existing cybersecurity and identity systems rather than replaces them, he added.
Petar Radanliev, AI security specialist at the University of Oxford’s Department of Computer Science in the U.K., said the risk goes beyond agents deliberately breaking through security controls.
“I think agents will go rogue. Not because anyone designs them to, but because they get confused,” he said. “The execution was excellent, and the judgment was poor, and in my view, execution is improving faster than judgment.”
In response, he termed Nvidia’s platform a “sensible answer” to the execution problem.
“If you cannot trust an agent's judgment, you limit what it can reach, and you watch it from hardware it cannot touch,” he said.
The safety platform
The system is made up of two main components: the open source-based OpenShell, which establishes boundaries on what an agent can access, and Sentry, a reference system design that runs on Nvidia’s BlueField-4 data processing units and provides a separate layer of enforcement outside the agent’s normal software environment.
According to Nvidia, Sentry can identify and quarantine an agent that attempts to move outside its permitted boundary in less than a second.”
The approach is intended to address a growing problem as companies move beyond AI assistants toward agents capable of taking actions across enterprise systems.
Writing on X, Nvidia CEO Jensen Huang said AI’s full promise can only be realized if users have confidence that it is built to be safe and deployed responsibly.
According to Nvidia, more than 100 organizations are working with it on the platform, including Anthropic, Cisco, CrowdStrike, Microsoft, Salesforce, SAP, Scale AI, ServiceNow, Palantir and Palo Alto Networks, as well as robotics companies including Figure, Gecko Robotics and Skild AI.
OpenShell, which works with open and proprietary AI systems, is now available through Nvidia’s developer resources and GitHub.
As for the Open Agent Safety platform, it enables safety enforcement to extend beyond the agentic environment, with OpenShell restricting access to credentials, networks, files and tools, and Sentry providing additional monitoring and enforcement.
Moreover, the fact that OpenShell is open and model-agnostic will attract developers and enterprises looking for a tighter monitoring system.
However, one challenge is that even if an agent is restricted and contained, it can still misbehave.
“An agent can stay within every technical boundary and still make a bad decision, misunderstand a task or misuse permissions it was legitimately given,” Kompella said.
Unclear if Nvidia approach is effective
Radanliev warned that the platform’s efficacy has not yet been demonstrated.
“Nobody, including Nvidia, can yet show that any of these platforms works,” he said. “There is no shared benchmark and no published comparative testing. Buyers are comparing architectures, not results.”
He also highlighted the question of what happens when an agent stays within those boundaries while making the wrong decision. An error or hallucination could, for example, be written into an agent’s memory or passed from one agent to another as fact and cause it to influence a decision weeks later without triggering a security alert, he said.
“Runtime monitoring catches an agent that breaks a rule. It is far less good at catching one that is confidently wrong,” Radanliev said.
For businesses deploying agents, this means security controls need to extend beyond monitoring what agents do.
“Don't assume a competent agent is a sensible one,” Radanliev said. “Keep a human decision on anything irreversible, because deciding is what agents are worst at.”
He also recommended auditing what agents remember and communicate with one another, as well as ensuring organizations can trace the source of errors.
“If you cannot trace an error back to its source, you cannot find the sleeper,” he added.
Scarlett Evans is a freelance writer with a focus on robotics, AI and emerging technologies. Previously, she was assistant editor at IoT World Today, where she specialized in robotics and smart city technologies. Scarlett also has a background in the mining and resources sector, with experience at Mine Australia, Mine Technology and Power Technology.
Esther Shittu is an Informa TechTarget news writer and podcast host covering AI technology.