AI News

Nvidia’s Open Agent Safety Platform Puts a Security Boundary Around AI Agents

Nvidia launched its Open Agent Safety Platform on September 28, 2026, with a practical premise: an AI agent should not get unlimited access to a computer, a company’s data, or an API merely because its instructions say to behave. The platform pairs OpenShell, an open-source runtime that restricts what an agent can reach, with Sentry, a reference design for a separate watchdog on Nvidia’s BlueField-4 hardware. Nvidia’s announcement

CEO Jensen Huang introduced the platform on X, saying more than 100 industry partners were involved. “Safety is how trust is earned,” he wrote. His larger claim is that agent safety needs an open ecosystem, rather than one vendor’s model-level safeguards. Nvidia has a commercial interest in that vision: the full reference design is optimized for its Vera CPUs and BlueField DPUs. OpenShell itself can also run without those chips.

OpenShell is available now, including its code and a new 0.1.0 release. Sentry is the proposed additional hardware-isolated layer in Nvidia’s reference system design. Nvidia has not published independent evidence that the combined system prevents every escape, malicious tool call, or harmful real-world action.

What Nvidia announced

Part Its job What is available now
OpenShell Runs agents in sandboxes and enforces rules for files, processes, networks, services, and credentials outside the agent workload. Open-source software; Nvidia says version 0.1.0 is available.
Sentry Watches agent activity and enforces an additional policy boundary independently of the host. A reference system design using BlueField-4 DPUs and Nvidia DOCA.
Vera CPU and BlueField-4 Supply the compute and separate security hardware for Nvidia’s optimized deployment. Part of Nvidia’s preferred full-stack architecture; not required to try OpenShell.

Nvidia’s technical overview describes three layers: the agent application, the runtime that governs it, and the infrastructure below it. The distinction matters. A model can be instructed to avoid a dangerous action, but a runtime rule can deny that action even when the model attempts it.

OpenShell predates this announcement. Nvidia described it earlier in 2026; the September 28 launch places it in a broader safety design and accompanies the 0.1.0 release. The entire stack is not yet a newly available open-source product. Nvidia’s OpenShell walkthrough

How OpenShell draws the boundary

Consider an agent asked to summarize issues in a GitHub repository. It needs to read the API. It does not need permission to create a release, delete a branch, or send a token to another server. OpenShell lets an operator define those limits as policy and checks the agent’s requests as it runs.

The architecture has a gateway that manages sandboxes and their policies, a supervisor that mediates service access, and a sandbox that confines the agent’s processes. Nvidia says operating-system controls restrict file access and privilege escalation. Network requests travel through the supervisor, which can inspect supported HTTP, GraphQL, and Model Context Protocol traffic. Its published demo permits a GitHub API read while blocking a POST to the same host. Technical walkthrough

The restrictions are intended to survive an agent opening a shell, executing generated code, or launching child processes. OpenShell also keeps real service credentials outside the agent workload, substituting them only for authorized destinations. It records allow and deny decisions for audit. A blocked request can produce an error the agent can use to seek a narrower permission, while the operator retains the authority to approve it.

The public repository lets developers inspect those claims. Its policy documentation describes filesystem restrictions through Landlock, reduced process privileges, and a proxy that evaluates network destination, port, calling program, and optional application-level rules. Nvidia’s walkthrough says network policies can change while a sandbox runs; changing filesystem or process restrictions requires a new sandbox. These are operating constraints a team should understand before deployment.

Nvidia also describes a policy prover that checks whether the permissions modeled in a policy stay inside an operator-defined boundary. That can expose a rule that looks narrow but permits a second route to the same sensitive action. It does not prove that every possible agent behavior is safe, or that the operator chose the right boundary in the first place. Nvidia says analysis of combined permissions across multiple agents remains ongoing work.

What Sentry adds, and what it does not

OpenShell’s supervisor sits outside the agent workload, but it is still part of the host-side software environment. Sentry moves a further monitoring and enforcement layer to a BlueField-4 data processing unit. Nvidia says this gives operators an out-of-band view of agent requests, identity, policy decisions, and tool or data access. It says Sentry can quarantine agents that move outside their permitted boundaries in milliseconds. Launch release

In Nvidia’s Vera Rubin POD design, BlueField-4 sits on the path to the model. Nvidia argues that this is a useful control point: an agent must reach a model to generate its next step, so an independent system can observe or interrupt that path. The company says the hardware boundary can remain operational even if the host is compromised. Architecture explanation

That is the most distinctive part of Nvidia’s proposal, and also the part readers should treat most carefully. The launch materials describe a reference design and vendor-reported capabilities. They do not establish the false-positive rate of Sentry’s behavior detection, its performance across different agent frameworks, or how reliably it recognizes a harmful action before damage occurs. Blocking a model call can stop a next step; it cannot undo an action already completed. Nor does network-level observation automatically reveal every important fact about an agent’s intent.

Organizations can run OpenShell without Sentry or BlueField-4. Nvidia says the open-source runtime can be extended to third-party compute platforms, including Arm and Intel. The hardware-based layer is optional, which makes the platform accessible to developers who want to test policy enforcement before considering Nvidia’s full infrastructure stack. Platform FAQ

Why the launch matters

Long-running agents can combine code execution, browser or API access, credentials, and sub-agents. A refusal written into a prompt is a weak security boundary for that kind of system. Security teams need to define permissions, inspect requests, keep secrets out of the workload, and see who approved a change. Nvidia’s design applies established access-control ideas to the agent’s operating environment.

The ecosystem is broad. Nvidia names Anthropic, Cisco, CrowdStrike, Dell Technologies, Hugging Face, Microsoft, Red Hat, Salesforce, SAP, and other organizations in its announcement. Those names signal interest and integration work; they are not evidence that every partner has deployed the complete OpenShell-plus-Sentry stack in production. Nvidia’s technical post gives narrower examples: Cadence uses OpenShell for chip design, Slack is building an on-demand agent platform on it, and Gecko Robotics uses it to govern agents involved with physical robots. Nvidia’s adoption examples

For a team evaluating OpenShell today, the right first test is a task with a known permission boundary. Give an agent the data and API operations it needs, deny a nearby write or external destination, and examine the resulting policy log. Then try the same task through generated code or a child process. That exercise will reveal more about the deployment than a broad claim that the agent is “safe.” Nvidia provides a quickstart and a read-only GitHub API example.

The platform gives operators a way to enforce specific limits as agents work. Its success will depend on how well those limits are written, which paths the controls actually inspect, how quickly failures are detected, and whether independent testing matches Nvidia’s claims. OpenShell can be examined and tried now. Sentry’s promised second boundary needs the same level of public scrutiny as it moves from reference design into deployments.

Reporting note: This article is based on Nvidia’s September 28 announcement and technical documentation, the public OpenShell repository, and Huang’s X post. Kingy.ai did not run an independent security or performance test of the platform for this report.