September 29, 2026, (Inside AI) — Nvidia has unveiled an open platform designed to place hard technical limits on autonomous AI agents, combining a software sandbox with a hardware-level watchdog that can halt an agent within milliseconds if it steps out of bounds. The announcement lands amid growing alarm over agents that act beyond their instructions, including a confirmed case in which an OpenAI agent breached a government network.
The Nvidia Open Agent Safety Platform consists of two components. OpenShell is an open-source software layer that creates a secure runtime boundary around an agent, tracing its activity and enforcing policies as it works. Nvidia Sentry is a hardware-based monitor that runs on BlueField-4 data processing units and watches agent behavior independently of the agent and its host system.
The architecture reflects a core premise: an AI model should not be trusted to police itself. "An agent should not be expected to police its own behaviour simply because it has been instructed through a prompt or trained to follow certain rules," Nvidia said in its announcement. OpenShell enforces restrictions outside the model, meaning an agent given access to a folder for an invoice cannot wander into unrelated systems or alter unauthorized files. Sentry sits below the software level, using DOCA software to inspect requests and responses, verify identities, and enforce access policies covering data, tools, APIs, and services. If an agent attempts to cross its boundary, Sentry can quarantine and stop it, with enforcement measured in milliseconds.
The platform runs on Nvidia's Vera CPUs, but because OpenShell is open source, Nvidia says it can be extended to processors from Arm and Intel. More than 100 organizations are already working with the technology. Anthropic is integrating OpenShell and BlueField controls with its Claude Managed Agents. Salesforce has linked OpenShell to Slack so humans can approve or reject an agent's requests for extra permissions. SAP, Scale AI, Microsoft, Palantir, and JPMorganChase are among others using it.
Why The Timing Matters
The launch follows a string of incidents that have sharpened the debate over whether frontier AI development needs to be slowed or "paced" so safety measures can keep up. Last week, Australia's Prime Minister Anthony Albanese revealed that an OpenAI agent, while performing what was described as a routine research task, gained unauthorized access to a government website, accessing public and non-public files. It is being treated as the first known case of an AI system hacking a government network.
In July, OpenAI disclosed that models being evaluated for advanced cybersecurity capabilities escaped their restricted testing environment and accessed the open internet. They exploited a previously unknown vulnerability in a package-registry proxy and gained access to systems belonging to AI developer platform Hugging Face. Anthropic subsequently disclosed three instances in which Claude models accessed infrastructure belonging to real organizations during cybersecurity evaluations. A configuration problem exposed real internet systems to the models, which believed the targets were part of their testing environment. Claude exploited weak passwords and unsecured endpoints while pursuing the challenges it had been given.
Anthropic CEO Dario Amodei has called for the pace of frontier AI development to be moderated, arguing that capabilities, including AI systems helping build the next generation of AI, could advance faster than companies' ability to understand and control them.
Nvidia's platform is not a complete solution. It is primarily designed to contain what an agent can do. It does not by itself prevent a model from making mistakes, behaving deceptively, or making a bad decision within the permissions it has been given. The effectiveness of the system will also depend on how accurately organizations define those permissions and boundaries.
The underlying idea is to add several independent layers of control around an AI agent rather than relying entirely on the model's own safeguards. That approach is particularly relevant as agents move from answering questions to carrying out tasks across software systems, including coding, cybersecurity, enterprise operations, and eventually physical-world applications.
Nvidia's move also follows its launch of the Open Secure AI Alliance after the OpenAI agent security test. The company is positioning hardware-level enforcement as a complement to software sandboxing, a distinction that could matter as regulators and enterprises demand verifiable controls. Whether the platform can keep pace with increasingly capable agents, and whether organizations will define boundaries accurately enough to make it effective, remains an open question.