According to Zhitong Finance APP, as debates around security risks in the artificial intelligence (AI) industry grow increasingly fierce, Nvidia (NVDA.US) has opted to address these concerns with engineering solutions. On Monday, Nvidia released a two-layer AI security platform called the "Open Agent Safety Platform." Through an architecture divided into application layer, runtime, and infrastructure, the platform provides continuous monitoring, real-time policy enforcement, and security governance for AI agents.
This system includes two open-source software tools: OpenShell and Nvidia Sentry. OpenShell runs agents in a sandbox environment, turns user instructions into verifiable safety policies, and restricts agent access to files, networks, tools, processes, and credentials. Nvidia Sentry extends monitoring and security controls to the BlueField data processing unit, using Nvidia DOCA for associating agent interactions, policy decisions, and tracking tool and data access. This allows for identifying abnormal behavior and intervening where necessary. In other words, OpenShell is responsible for controlling the resources and execution privileges that AI agents can access, while Nvidia Sentry is responsible for monitoring agent behavior and identifying potential threats.
Nvidia stated that AI agents may deviate from set tasks due to factors such as tools, runtime, or ambiguous instructions, so independent security controls are necessary. The system is optimized for Nvidia Vera CPUs and BlueField DPUs and is compatible with other hardware systems. Nvidia also noted that it is working with Arm (ARM.US) and Intel (INTC.US) to ensure compatibility with these companies' CPUs as well.
Ali Golshan, Nvidia’s Senior Director of AI Software, said that these tools use mathematical formulas to detect whether AI agents are attempting to circumvent restrictions. For example, an agent could try to “spawn” multiple “sub-agents” to bypass blocks placed on the main agent. He commented: “What we’re really discussing here is agentic behavior, which is a cluster composed of multiple agents and how they operate collaboratively.”
Nvidia further indicated that the system is able to isolate anomalous agents within milliseconds, and claimed that if relevant labs had previously deployed this technology, the July OpenAI model attack on Hugging Face could have been prevented.
Justin Boitano, Vice President and General Manager of Nvidia’s Enterprise Computing Department, said: “Based on what we know today, if this brand-new security platform had been used during the initial model evaluation phase in cutting-edge AI labs, it could have stopped this breach.” He added: “We are advancing this effort openly and hope to collaborate with everyone and encourage joint participation.”
Nvidia released this latest AI security system at a time when security incidents related to AI agents continue to attract industry attention, most notably the July OpenAI model’s uncontrolled invasion of the globally renowned open-source AI platform Hugging Face. After the Hugging Face breach, several other security incidents involving OpenAI AI agents were revealed. For example, in September, it was reported that an OpenAI agent had taken control of a long-unmaintained German wiki site earlier this year. Furthermore, OpenAI stated in a mid-September blog post that six "unexpected or concerning model behaviors" had been discovered over the previous six months. These newly disclosed AI security incidents include concealing errors, seeking unauthorized credentials, uploading files to public websites, and facilitating communication between otherwise isolated training environments, with the earliest incident dating back to October of last year.
With frequent AI security incidents occurring recently, AI companies are under increased pressure to take security risks more seriously. On September 12, Anthropic CEO Dario Amodei published a cautionary article calling for a slowdown in frontier AI model development speeds. This call was quickly echoed by both Musk and Altman, two key figures in the industry.
However, as AI industry giants unusually call in unison to "tap the brakes," Nvidia CEO Jensen Huang has taken a strong stance against Anthropic’s call to "slow down" AI development. Huang does not oppose AI safety itself; rather, he is against drawing industry-wide slowdowns or halts directly from unverified catastrophic predictions. In contrast, Huang would prefer to treat AI safety as an engineering challenge. If an incident occurs in a leading lab, the primary response should be root cause analysis: What exactly happened? Which mechanisms failed? What measures could have been taken? What new technologies and processes should be established? Through sandboxing, continuous monitoring, verification, and evaluation, the aim is to prevent recurring issues before the next release.
According to Huang’s perspective, one of the main reasons today’s cutting-edge labs are more prone to exposing risks is because they wield the most computing power and tackle the hardest problems. As these companies transition from research institutes into true engineering organizations, they likewise need to establish corresponding testing and control systems.
From the launch of this AI security system, it is evident that Nvidia is aiming to further embed AI safety capabilities into the core infrastructure for model operation and agent execution. By implementing privilege controls and anomaly isolation, Nvidia is seeking to reduce the security risks posed by AI agents. This also suggests that Nvidia’s business is extending beyond AI chips into software and security infrastructure.