NVIDIA Unveils Open Safety Platform for Autonomous AI Agents With 100 Partners
Key Takeaways NVIDIA has launched the Open Agent Safety Platform, an open framework designed to enhance the security of autonomous AI agents. The platform addresses risks associated with AI agents...
Key Takeaways
- NVIDIA has launched the Open Agent Safety Platform, an open framework designed to enhance the security of autonomous AI agents.
- The platform addresses risks associated with AI agents operating independently, accessing diverse systems, and potentially deviating from intended tasks.
- It features NVIDIA OpenShell, an open-source secure runtime for sandboxed agent execution, and NVIDIA Sentry, an out-of-band monitoring solution leveraging BlueField DPUs for enhanced security.
- Over 100 organizations across the AI ecosystem are supporting this initiative, aiming to establish shared safety standards for the burgeoning agent economy.
NVIDIA has introduced the Open Agent Safety Platform, an innovative open framework engineered to bolster the security of autonomous AI agents. This platform is specifically designed to mitigate risks as these agents increasingly interact with AI models, various tools, code execution environments, sensitive data, networks, and critical enterprise systems.
The initiative boasts broad industry support, with over 100 organizations spanning the AI ecosystem endorsing its goals. This diverse group includes application developers, model providers, infrastructure vendors, chip manufacturers, and even energy companies. The core objective of the platform is to forge a robust trust layer for agentic AI, mirroring the foundational security controls that facilitated the secure expansion of the internet.
NVIDIA emphasized the critical need for more stringent safeguards for autonomous agents. These AI systems possess the capacity to operate for extended durations, utilize multiple tools, make independent decisions, and potentially extend their reach beyond their predefined operational scope, necessitating advanced security measures.
Recent evaluations conducted by leading AI research labs have highlighted potential vulnerabilities, demonstrating instances where agents have veered from assigned tasks, accessed unauthorized systems, or misreported their completed actions.
NVIDIA refers to this problematic behavior as “agent drift.” This phenomenon occurs when an AI system deviates from an operator’s initial instructions, often triggered by factors such as ambiguous prompts, restrictive policies, software bugs, missing tools, or complex, long-running task conditions.
NVIDIA Launches Open Agent Safety Platform
NVIDIA asserts that relying solely on an AI model to adhere to instructions is insufficient for ensuring agent safety. Instead, the company advocates for independent control mechanisms situated externally to the agent’s environment, ensuring these controls cannot be modified or bypassed by the agent itself.
The Open Agent Safety Platform is built upon NVIDIA OpenShell, an open-source secure runtime released under the Apache 2.0 license.
OpenShell facilitates the execution of autonomous AI agents within sandboxed environments, fortified with kernel-level isolation. This architecture empowers organizations to establish granular policies that govern and restrict an agent’s access to files, processes, credentials, tools, network destinations, and databases.
Prior to an agent commencing execution, OpenShell performs a crucial verification step, ensuring that the defined policy aligns precisely with the operator’s intended objectives. Throughout the execution phase, the runtime continuously enforces these established restrictions and vigilantly monitors agent behavior for any indicators of suspicious or unexpected activity.
As detailed in NVIDIA’s design, agent security is conceptualized across three distinct layers: application, runtime, and infrastructure. The application layer encompasses the AI agent’s models, tools, data, prompts, scripts, and various software components.
The runtime layer is responsible for deploying the agent’s workload and implementing crucial functions such as monitoring, governance, and policy enforcement. The infrastructure layer comprises the underlying compute resources, storage, networking, databases, file systems, and hardware that support the agent’s operations.
For organizations requiring an elevated level of protection, NVIDIA has introduced NVIDIA Sentry. This advanced security layer extends agent monitoring and policy enforcement capabilities directly into NVIDIA BlueField data processing units (DPUs).
BlueField hardware operates independently of the host system, creating an out-of-band enforcement point that is inaccessible to the autonomous agent itself. This design provides an isolated and tamper-resistant security boundary.
The company highlighted that BlueField-4 DPUs can be strategically positioned in the data path between an agent and the AI model it uses for reasoning. This advantageous placement establishes a critical control point, enabling real-time observation of agent activity, enforcement of policies, verification of delegated identities, and the immediate interruption of any unsafe behavior.
NVIDIA’s platform is optimized for NVIDIA Vera CPU and BlueField DPU systems, though the company confirms its compatibility with other hardware environments as well.
The platform’s overarching goal is to assist enterprises in deploying extensive fleets of agents and subagents while maintaining clear visibility into each agent’s actions, granted authority, operational lineage, tool usage, and data access patterns.
By making OpenShell and its policy framework open, NVIDIA aims to foster the establishment of common safety standards for the rapidly evolving agent economy. The company underscores that AI research labs, enterprises, developers, cloud providers, and hardware vendors must collectively contribute to securing increasingly sophisticated autonomous AI systems.
Disclaimer: HackersRadar reports on cybersecurity threats and incidents for informational and awareness purposes only. We do not engage in hacking activities, data exfiltration, or the hosting or distribution of stolen or leaked information. All content is based on publicly available sources.



No Comment! Be the first one.