NVIDIA Moved the Agent Watchdog to a Different Chip

The Open Agent Safety Platform puts enforcement on a DPU, outside the host the agent runs on. The instinct is decades old. The shape is the question.


Out-of-band supervision: an in-band guardrail runs inside the same host and trust domain as the agent, so a compromise of the host disables it; an out-of-band watchdog on a separate device observes and enforces across a hardware boundary the host cannot cross

Look at where your agent guardrails run.

The policy check lives in the orchestrator process. The tool allowlist is a config file the agent’s own runtime reads. The approval gate is application code. The audit log is written by the same service doing the acting, to a destination that service has credentials for.

Every one of those is inside the thing it’s supposed to constrain. They share a process, a host, a network namespace, a credential set. Whatever compromises the agent compromises its supervisor at the same instant, and the supervisor’s last act is to write a log entry saying everything looked fine.

That’s the problem NVIDIA’s Open Agent Safety Platform, announced on 28 September, is aiming at. I think the diagnosis is correct. I have more questions about the prescription.

What they actually shipped

Two pieces, at two levels.

OpenShell is runtime software that sets execution boundaries for agents on CPUs — sandboxing, policy enforcement, tracing what the agent does. It’s open source, NVIDIA says it’s broadly available now, it runs on their Vera CPUs, and they’re explicit that it can be extended to third-party platforms including Arm and Intel. That last detail matters more than it sounds; a safety runtime that only works on one vendor’s chips is a product, not a standard.

Sentry is the interesting one. It’s described as an out-of-band watchdog running on BlueField-4 DPUs. It watches agent behaviour continuously, inspects requests and responses, applies zero-trust policy to data, tools, APIs and services, and — the headline claim — quarantines an agent in milliseconds if it crosses its boundary. The enforcement happens in silicon, on a device that is not the machine the agent is running on.

Both are framed as open-source contributions to the Open Secure AI Alliance, under Linux Foundation governance. There are over a hundred named partners, including Anthropic, Microsoft, Cisco, CrowdStrike, Dell, HPE, Palantir, Salesforce, SAP and SpaceX AI.

That partner list is doing real work in the announcement, and it’s worth separating two things: a hundred logos indicates the industry agrees the problem is real. It does not indicate a hundred production deployments.

This instinct is much older than agents

Here’s why I take the architecture seriously despite the marketing around it: it isn’t new. It’s one of the oldest patterns in systems engineering, applied to a new subject.

A hardware watchdog timer is a counter the software must periodically reset. If the software wedges, the counter expires and the hardware reboots the machine. The entire point is that the thing being watched cannot stop the watcher — software cannot talk a watchdog out of firing by being confidently wrong.

A baseboard management controller — your iLO, your iDRAC — is a small computer bolted onto a big computer, with its own processor, its own network interface and its own power domain. It can power-cycle the host, mount media and read sensors while the host OS is dead or lying. It exists because management that depends on the managed system working is not management.

A hardware security module holds keys the host can use but never read. The security property comes from a boundary the host physically cannot cross, not from the host behaving.

Smart NICs and DPUs already enforce network policy below the hypervisor, so a compromised tenant can’t rewrite its own firewall rules.

In every case the reasoning is identical: a control that shares a fate with the thing it controls is not a control. It’s a suggestion with good intentions.

Applying that to agents is genuinely new. The underlying idea has fifty years of production behind it, and that’s a reason to expect the general shape to hold even if this specific product doesn’t.

What it fixes, stated narrowly

Be precise about the threat model, because the marketing language around “AI safety” blurs several very different problems together.

Out-of-band enforcement helps when the agent’s host is compromised or the agent’s own reasoning is subverted. Prompt injection convinces the agent to exfiltrate data; the agent is doing exactly what an attacker asked; the in-process policy check is evaluating rules the attacker now influences. A watcher on separate hardware sees the egress attempt regardless of what the agent believes it’s doing, because it isn’t reading the agent’s intentions — it’s reading the traffic.

It also gives you an audit trail the agent cannot edit. If the record of what an agent did is written by the agent’s own runtime, it’s evidence only as long as nothing went wrong, which is precisely when you need it. I’ve made this argument about authorization being an audit-trail problem and this is the hardware answer to it.

And it addresses the approval-fatigue failure mode from a different angle. If enforcement is mechanical and happens below the application, you don’t need a human clicking through gates to maintain the invariant — which is good, because that control degrades exactly as the volume it governs rises.

What it does not fix: an agent that does something harmful entirely within its permitted boundary. If the agent is allowed to call the deletion API and it calls the deletion API for the wrong reason, Sentry sees a permitted call. Policy enforcement answers “was this allowed,” never “was this wise.” Every control at this layer has that limit, and no amount of silicon changes it.

What I’d want to know before believing the pitch

Four things, none of which are answered by a press release.

What can it actually see? Agent traffic is TLS. A watcher on the wire either terminates TLS — becoming a decryption point, with everything that implies for key management and for its own attractiveness as a target — or it reasons about metadata: destinations, volumes, timing, connection patterns. Both are defensible. They’re very different products with very different detection ceilings, and “inspect agent requests and responses” doesn’t tell me which one this is.

Who writes the policy, and in what language? This is where every system of this kind actually lives or dies. A hardware enforcement point with an unusable policy language ends up configured permissively, because the alternative is breaking production. The hard part of policy is never the enforcement — it’s expressing intent precisely enough to enforce, and keeping it current as the application changes. Nothing about a DPU makes that easier.

What is the overhead, and what’s the failure mode? NVIDIA says “minimal overhead” for OpenShell, which is a phrase that means nothing without numbers. More importantly: what happens when Sentry is unavailable? Fail open and the control is advisory, which is the thing we’re trying to escape. Fail closed and your DPU is now a single point of failure for every agent on the host. That’s a legitimate engineering choice; I’d just like it stated.

Does it survive heterogeneity? Most organisations run agents in several places — a cloud provider’s managed service, somebody’s SaaS, a laptop. A control that only works on hardware you own covers the fraction of your fleet you were always most able to secure, and misses the rest. OpenShell being portable to Arm and Intel is the right instinct. Sentry, by construction, is not portable.

The part that doesn’t need a DPU

Here’s where I’d push back on how this will get read.

The valuable idea in this announcement is architectural, not physical. It’s put the enforcement point outside the agent’s blast radius. Separate silicon is one implementation — the strongest available, and the most expensive, and the least applicable to where most teams actually are.

You can implement the same principle today, with what you have:

Make tool access a call to a broker, not a library the agent imports. The moment the agent’s code can reach the implementation directly, the policy check is advisory. This is the whole argument for the MCP gateway pattern, and it’s the single highest-leverage move on this list.

Issue the agent credentials scoped to one task, from a service it doesn’t control, so its capabilities come from outside rather than from a config file it can read and an attacker can influence. Not a shared service account — identities, not API keys.

Write the audit record from the broker, to a store the agent has no write path to. If the agent can edit the log, you don’t have a log.

Evaluate policy in a service with its own deployment lifecycle, so changing what an agent may do isn’t a code change in the agent.

None of that requires new hardware. All of it moves enforcement outside the process being enforced, which is the actual property. Teams that have done those four things get most of the benefit; teams that haven’t will not be rescued by buying a DPU, because their gap is architectural and they’d be adding a hardware control to a system whose software controls are still in-band.

Which is the thing I’d actually say to a platform lead this week: the announcement is a useful prompt to go and look at where your enforcement runs. Draw the trust boundary on the diagram. Mark every guardrail that lives inside it. If most of them do — and they will — you’ve found work that is cheaper than procurement and available now.

The silicon question can wait until the architecture question is answered. It’ll still be there, and by then there may be more than one vendor answering it, which is a better week to be buying.


Related: Identity was the easy half — agent authorization is an audit-trail problem · The MCP gateway pattern · The 47th approval click is the bug

Frequently asked questions

What is NVIDIA's Open Agent Safety Platform?

Announced on 28 September 2026, it pairs two components. OpenShell is open-source runtime software that sets execution boundaries for agents running on CPUs; NVIDIA describes it as broadly available, running on its Vera CPUs and extensible to third-party platforms including Arm and Intel. Sentry is an out-of-band watchdog running on BlueField-4 DPUs that monitors agent behaviour continuously and enforces policy in silicon, and which NVIDIA says can quarantine a misbehaving agent in milliseconds. Both are presented as open-source contributions under the Linux Foundation-governed Open Secure AI Alliance.

Why does it matter that enforcement runs on separate hardware?

Because a control that runs on the same host, in the same trust domain, as the thing it constrains can be disabled by whatever compromises that host. Almost every agent guardrail in common use today — a policy check in the orchestrator, an allowlist in the tool layer, an approval gate in the application — shares a fate with the process it is supposed to govern. Moving enforcement to a separate device means the compromise has to cross a hardware boundary before the watcher can be silenced.

Is out-of-band supervision a new idea?

No, and that is a point in its favour. Hardware watchdog timers that reset a wedged machine, baseboard management controllers that operate independently of the host OS, hardware security modules that hold keys the host can never read, and smart NICs that enforce network policy below the hypervisor are all the same instinct: put the control somewhere the controlled thing cannot reach. Applying it to AI agents is new; the architecture is decades old and well understood.

Do I need a DPU to supervise my agents?

Almost certainly not yet. The principle — put the enforcement point outside the agent's blast radius — is implementable today with ordinary infrastructure: a broker the agent must call rather than library code it imports, credentials issued per task, policy evaluated in a service the agent cannot modify, and an audit log it cannot write to. Most organisations have not done those things. Buying specialised silicon before doing them is building the second floor first.

Comments