
NVIDIA’s Open Agent Safety Platform Puts a Control Layer Between Agents and Power
NVIDIA’s OpenShell and Sentry reference design treats autonomous agents like untrusted software, combining sandboxing, DPU monitoring, policy enforcement, and audit trails.
NVIDIA’s Open Agent Safety Platform Puts a Control Layer Between Agents and Power
NVIDIA’s new safety platform is a hardware-and-runtime argument: agents should not be trusted merely because a model provider says they are aligned. The proposed stack puts policy enforcement outside the agent, then watches the path between the agent, its tools, and the model.
The primary source is NVIDIA's September 28, 2026 reference design. Product architecture and platform claims below are attributed to NVIDIA; the practical analysis separates those claims from controls an operator would still need to test.
flowchart LR
A[User task] --> B[Agent plan]
B --> C[Article-specific evidence or tool boundary]
C --> D[Verification and policy]
D --> E[Human or controlled outcome]
Agents need a browser-like trust boundary
A web page can be useful and untrusted at the same time. The browser made that combination workable by limiting what page code could do. NVIDIA applies the same intuition to agents: a model can plan and act, but the runtime should assume the plan may drift, the tool may be unsafe, or the instruction may be ambiguous.
The release record gives this discussion a concrete anchor: NVIDIA published the Open Agent Safety Platform reference article on September 28, 2026. NVIDIA lists verifiable policy, out-of-band enforcement, model-path control, authority scaling, and shared responsibility as core principles. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
OpenShell is the software perimeter
OpenShell supplies sandboxing and kernel-level isolation around the agent. That matters because prompt rules are not a security boundary. A model can misunderstand a rule, a tool can return hostile content, and a dependency can be compromised. The runtime needs to constrain filesystem, network, process, and credential access regardless of what the model says.
The release record gives this discussion a concrete anchor: OpenShell is described as an Apache 2.0 open-source runtime for agents in sandboxed environments with kernel-level isolation. Sentry is described as extending monitoring and enforcement into BlueField hardware through NVIDIA DOCA. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
Sentry moves observation below the agent
A model can misreport its own actions. NVIDIA’s design responds by putting Sentry on BlueField-4 and using DOCA to observe and enforce around the workload. Out-of-band monitoring is valuable because the component being monitored cannot rewrite the record of what crossed the enforcement point. It also makes hardware placement and firmware trust part of the AI safety conversation.
The release record gives this discussion a concrete anchor: The reference design combines OpenShell on NVIDIA Vera CPUs with NVIDIA Sentry on BlueField-4 DPUs. The reference design places BlueField-4 DPUs on the node path to the model in NVIDIA Vera Rubin POD systems. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
Policy must be verifiable
A policy that exists only in a prompt is difficult to test. Verifiable policy means the system can state which rule applied, which identity requested the action, and whether the requested scope matched the rule. That creates a decision record that can be inspected after an incident instead of relying on a transcript written by the same agent that caused the incident.
The release record gives this discussion a concrete anchor: NVIDIA lists verifiable policy, out-of-band enforcement, model-path control, authority scaling, and shared responsibility as core principles. NVIDIA defines drift as agent actions departing from intended task or operating constraints. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
The model path is an authority boundary
If an agent can reach a model, a tool, and a data store through separate routes, defenders may miss the relationship between them. NVIDIA’s reference design emphasizes controlling the path to the model and correlating interactions. The practical goal is to see that a model requested a tool, the tool returned data, and the next action crossed a policy boundary.
The release record gives this discussion a concrete anchor: Sentry is described as extending monitoring and enforcement into BlueField hardware through NVIDIA DOCA. The platform is presented as monitoring agent interactions, policy decisions, and tool access in contextual activity records. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
Drift is more common than malice
An agent may drift because a page changed, a tool failed, a policy blocked its first plan, or an instruction was incomplete. Treating every deviation as malicious makes systems unusable; ignoring it makes them unsafe. The runtime needs thresholds, explainable interventions, and a way to pause for a human when the agent’s plan no longer matches the task.
The release record gives this discussion a concrete anchor: The reference design places BlueField-4 DPUs on the node path to the model in NVIDIA Vera Rubin POD systems. NVIDIA separately published a guide for adding runtime controls to agents with OpenShell. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
Authority should scale with visibility
NVIDIA proposes scaling agent authority with reasoning visibility. The operational translation is straightforward: an agent with a clear plan, narrow tools, and observable intermediate state can receive more autonomy than an agent whose actions cannot be explained or replayed. Visibility is not a guarantee of correctness, but it gives operators a basis for deciding how much power to grant.
The release record gives this discussion a concrete anchor: NVIDIA defines drift as agent actions departing from intended task or operating constraints. The article compares the needed trust layer to browser sandboxing for web software. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
Hardware enforcement has a price
A DPU-based layer can inspect traffic and enforce policy without consuming the agent’s own compute path, but it adds deployment complexity. Teams must manage firmware, device compatibility, update channels, telemetry volume, and failure behavior. Hardware enforcement is most useful where the workload already has a standard infrastructure layer; it is harder to justify for a small local agent.
The release record gives this discussion a concrete anchor: The platform is presented as monitoring agent interactions, policy decisions, and tool access in contextual activity records. NVIDIA published the Open Agent Safety Platform reference article on September 28, 2026. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
Why the Apache license matters
OpenShell under Apache 2.0 invites adoption and inspection, but open source does not equal secure by default. Enterprises still need to review the code, pin releases, harden images, and define who owns patches. The license improves the chance that the runtime becomes a shared control surface rather than a proprietary feature hidden inside one vendor’s model stack.
The release record gives this discussion a concrete anchor: NVIDIA separately published a guide for adding runtime controls to agents with OpenShell. OpenShell is described as an Apache 2.0 open-source runtime for agents in sandboxed environments with kernel-level isolation. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
The activity record must be useful
A log that says “tool called” is not enough. Operators need the agent identity, task identity, policy version, tool arguments after redaction, returned data classification, model request, enforcement decision, and outcome. Records should support incident reconstruction without becoming a second sensitive data lake. The platform’s contextual-record idea is sound only if retention and access are designed alongside capture.
The release record gives this discussion a concrete anchor: The article compares the needed trust layer to browser sandboxing for web software. The reference design combines OpenShell on NVIDIA Vera CPUs with NVIDIA Sentry on BlueField-4 DPUs. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
MCP and APIs need different controls
An MCP server may expose a dynamic tool description, while an API often has stable schemas and explicit credentials. Both need allowlists and validation, but the controls should reflect their failure modes. The runtime can require read-only MCP tools by default, constrain API methods, and block a tool response from silently expanding the agent’s authority.
The release record gives this discussion a concrete anchor: NVIDIA published the Open Agent Safety Platform reference article on September 28, 2026. NVIDIA lists verifiable policy, out-of-band enforcement, model-path control, authority scaling, and shared responsibility as core principles. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
A safer deployment pattern
A first pilot should run OpenShell with synthetic data, no production credentials, and a policy that denies network egress unless a named tool is approved. Operators can compare the model’s intended action with the enforced action, then introduce real read access. Write capabilities should arrive last, with idempotency keys and human approval for external side effects.
The release record gives this discussion a concrete anchor: OpenShell is described as an Apache 2.0 open-source runtime for agents in sandboxed environments with kernel-level isolation. Sentry is described as extending monitoring and enforcement into BlueField hardware through NVIDIA DOCA. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
Where the reference design is strongest
The proposal is strongest when it refuses to make the model the sole security layer. Sandboxing, hardware observation, and policy logs address different failure points. The design also gives infrastructure teams a place to participate in agent safety instead of delegating every control to a model lab.
The release record gives this discussion a concrete anchor: The reference design combines OpenShell on NVIDIA Vera CPUs with NVIDIA Sentry on BlueField-4 DPUs. The reference design places BlueField-4 DPUs on the node path to the model in NVIDIA Vera Rubin POD systems. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
What it cannot guarantee
A control layer cannot decide whether a business action is wise if the business policy is missing. It can block an unauthorized transfer, but it cannot know whether an authorized transfer is fraudulent without context. It also cannot repair bad data or an incorrect objective. The platform reduces blast radius; it does not replace governance, testing, or human accountability.
The release record gives this discussion a concrete anchor: NVIDIA lists verifiable policy, out-of-band enforcement, model-path control, authority scaling, and shared responsibility as core principles. NVIDIA defines drift as agent actions departing from intended task or operating constraints. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
The market implication
As agents gain access to production systems, runtime safety will become a purchasing criterion. Buyers will ask where policies execute, who can alter them, whether actions are independently recorded, and how the system behaves when the model is unavailable. NVIDIA’s platform is an early attempt to make those answers architectural rather than aspirational.
The release record gives this discussion a concrete anchor: Sentry is described as extending monitoring and enforcement into BlueField hardware through NVIDIA DOCA. The platform is presented as monitoring agent interactions, policy decisions, and tool access in contextual activity records. Those details are not decoration. They define the boundary of the claim and show where an implementation team would need to look before copying the idea into production.
For a platform team, the design only helps when the enforcement point survives the failure it is meant to contain. Test a compromised tool, a confused agent, a missing policy, a stale credential, a DPU outage, and a rollback. The system should fail closed for high-impact actions and fail visibly for low-impact work. Silent fallback to an unrestricted model path would erase the value of the control layer.
The test that will separate safety from theater
The useful demonstration is not an agent completing a clean task behind a polished dashboard. It is an agent that receives a malicious document, loses its network route, requests an unapproved credential, and then tries to explain what happened. The operator should see the attempted action, the policy that stopped it, the component that enforced the stop, and the recovery state. If the record depends on the agent's own story, the platform has not yet demonstrated independent control.
That test should run across software and hardware revisions. A policy can be correct in OpenShell while a deployment image, DPU firmware, or telemetry pipeline changes the result. Reproducible fixtures and signed policy versions give teams a way to compare releases. NVIDIA's reference architecture is most valuable if it encourages this kind of adversarial operations practice rather than becoming another vendor diagram displayed only during procurement.
The sources behind the story
The primary announcement and related technical references are listed below. Publication dates are kept distinct from the dates of the underlying work; vendor benchmark and performance claims are attributed to the organizations that published them.
- https://developer.nvidia.com/blog/nvidia-open-agent-safety-platform-a-reference-for-continuous-in-silicon-agent-monitoring/
- https://developer.nvidia.com/blog/add-runtime-controls-to-ai-agents-with-nvidia-openshell/
- https://github.com/NVIDIA/OpenShell
- https://www.nvidia.com/en-us/networking/products/data-processing-unit/
- https://developer.nvidia.com/networking/doca
- https://www.kernel.org/doc/html/latest/admin-guide/LSM/index.html
- https://www.nist.gov/itl/ai-risk-management-framework
- https://owasp.org/www-project-top-10-for-large-language-model-applications/
- https://modelcontextprotocol.io/specification/latest
- https://www.cisa.gov/resources-tools/resources/secure-by-design
The practical takeaway is narrow but durable: the useful AI system is the one whose evidence, authority, and failure boundary remain visible after the demo ends.