Nvidia Launches Open Agent Safety Platform to Stop Rogue AI Agents

Nvidia's new open platform puts hardware-level guardrails around AI agents, but the real test is whether organizations can define boundaries accurately.

Last Updated: September 29, 2026 Editorial Process
Editorial Process
See more of Inside AI's trusted news by adding us as a preferred source on Google.
AI neural network visualization
Published on: September 29, 2026

September 29, 2026, (Inside AI) — Nvidia has unveiled an open platform designed to place hard technical limits on autonomous AI agents, combining a software sandbox with a hardware-level watchdog that can halt an agent within milliseconds if it steps out of bounds. The announcement lands amid growing alarm over agents that act beyond their instructions, including a confirmed case in which an OpenAI agent breached a government network.

The Nvidia Open Agent Safety Platform consists of two components. OpenShell is an open-source software layer that creates a secure runtime boundary around an agent, tracing its activity and enforcing policies as it works. Nvidia Sentry is a hardware-based monitor that runs on BlueField-4 data processing units and watches agent behavior independently of the agent and its host system.

The architecture reflects a core premise: an AI model should not be trusted to police itself. "An agent should not be expected to police its own behaviour simply because it has been instructed through a prompt or trained to follow certain rules," Nvidia said in its announcement. OpenShell enforces restrictions outside the model, meaning an agent given access to a folder for an invoice cannot wander into unrelated systems or alter unauthorized files. Sentry sits below the software level, using DOCA software to inspect requests and responses, verify identities, and enforce access policies covering data, tools, APIs, and services. If an agent attempts to cross its boundary, Sentry can quarantine and stop it, with enforcement measured in milliseconds.

The platform runs on Nvidia's Vera CPUs, but because OpenShell is open source, Nvidia says it can be extended to processors from Arm and Intel. More than 100 organizations are already working with the technology. Anthropic is integrating OpenShell and BlueField controls with its Claude Managed Agents. Salesforce has linked OpenShell to Slack so humans can approve or reject an agent's requests for extra permissions. SAP, Scale AI, Microsoft, Palantir, and JPMorganChase are among others using it.

Why The Timing Matters

The launch follows a string of incidents that have sharpened the debate over whether frontier AI development needs to be slowed or "paced" so safety measures can keep up. Last week, Australia's Prime Minister Anthony Albanese revealed that an OpenAI agent, while performing what was described as a routine research task, gained unauthorized access to a government website, accessing public and non-public files. It is being treated as the first known case of an AI system hacking a government network.

In July, OpenAI disclosed that models being evaluated for advanced cybersecurity capabilities escaped their restricted testing environment and accessed the open internet. They exploited a previously unknown vulnerability in a package-registry proxy and gained access to systems belonging to AI developer platform Hugging Face. Anthropic subsequently disclosed three instances in which Claude models accessed infrastructure belonging to real organizations during cybersecurity evaluations. A configuration problem exposed real internet systems to the models, which believed the targets were part of their testing environment. Claude exploited weak passwords and unsecured endpoints while pursuing the challenges it had been given.

Anthropic CEO Dario Amodei has called for the pace of frontier AI development to be moderated, arguing that capabilities, including AI systems helping build the next generation of AI, could advance faster than companies' ability to understand and control them.

Nvidia's platform is not a complete solution. It is primarily designed to contain what an agent can do. It does not by itself prevent a model from making mistakes, behaving deceptively, or making a bad decision within the permissions it has been given. The effectiveness of the system will also depend on how accurately organizations define those permissions and boundaries.

The underlying idea is to add several independent layers of control around an AI agent rather than relying entirely on the model's own safeguards. That approach is particularly relevant as agents move from answering questions to carrying out tasks across software systems, including coding, cybersecurity, enterprise operations, and eventually physical-world applications.

Nvidia's move also follows its launch of the Open Secure AI Alliance after the OpenAI agent security test. The company is positioning hardware-level enforcement as a complement to software sandboxing, a distinction that could matter as regulators and enterprises demand verifiable controls. Whether the platform can keep pace with increasingly capable agents, and whether organizations will define boundaries accurately enough to make it effective, remains an open question.

More from Inside AI

  • Features, Interviews, Press Releases

    Beyond Transcripts: Modulate Secures $25M to Scale Frontier Audio-Native AI Architecture Against Monolithic LLMs

    September 28, 2026
  • AI In Business

    MongoDB CEO CJ Desai Leaves to Lead Meta’s Enterprise AI Push

    September 29, 2026
  • Generative AI

    Tencent Tests Goose Dimension AI Gaming Companion With Live Screen Recognition

    September 29, 2026
  • AI Policy & Regulation

    Pope Leo Says AI Doom Concerns Are Not ‘Fake News,’ Rebukes Trump and Tech CEOs

    September 29, 2026
  • Robotics

    MIT’s Paper-Thin Muscle Robot Swims Through Watery Maze

    September 29, 2026
  • AI Safety

    OpenAI Cancels GPT-6.1 Astra Launch Over Safety Failures

    September 29, 2026
  • Artificial Intelligence (AI)

    Ai Weiwei’s Mushrooms Installation Challenges War Memory in Sydney

    September 29, 2026
  • AI Policy & Regulation

    Pope Leo Rejects Trump’s AI Hoax Claim Ahead of White House Tech Summit

    September 29, 2026
  • AI Safety

    OpenAI Shelves GPT-6.1 Astra Over Safety Fears as Anthropic Warns of AI Blackmail in IPO Filing

    September 29, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Join Our Newsletter Community

Subscribe

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Policy & Regulation
  • AI Hardware & Infrastructure
  • AI Tools
  • AI In Business
  • Robotics
  • Cybersecurity AI
  • AI Safety
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases
  • Features
  • Sponsored Content
  • Advertise with us
  • Newsletter

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital