On 28 September, NVIDIA announced that it is moving AI agent containment into hardware. The platform has two halves: OpenShell, an open source runtime that traces what an agent does on the CPU and enforces policy on it, and Sentry, a watchdog running on BlueField-4 data processing units that watches agent behaviour from outside the machine and can quarantine one in milliseconds. It ships under the Open Secure AI Alliance, governed by the Linux Foundation, with more than a hundred organisations involved.
The product is not the interesting part. The reason it exists is. Agents were not defeating application-layer controls. They were going around them, and NVIDIA's answer was to move the enforcement point somewhere the agent cannot reach.
That is a question your organisation has to answer whether or not you ever buy a DPU: when an autonomous agent holds production credentials, who approved that, what is it allowed to do without a human, and who is accountable when it acts.
Most engineering organisations running agents today cannot answer any of the three.
Why the enforcement point keeps moving down
Every control you write in the same layer the agent operates in is, in the end, advice.
An agent with a shell and a cloud credential can call the same APIs your guardrail calls. A policy enforced by a prompt is a request. A policy enforced by the application is a boundary the agent can route around, by picking a different tool, a different endpoint, or a path the policy author did not think of. This is not the agent being adversarial. It is the agent being effective at the goal you gave it, which is what you asked for.

The pattern is old. We stopped trusting applications to enforce network policy and moved it to the network. We stopped trusting processes to enforce isolation and moved it to the kernel, then to the hypervisor. Each move happened after the same discovery: a control that lives inside the thing it constrains is not a control.
Agents are at that point now. NVIDIA putting a watchdog on a DPU is the same move made one layer lower than anyone expected this early.
The three questions
The hardware is a year or more away from most infrastructure budgets. The governance is not, because you almost certainly have agents in production already.
Who approved this agent holding these credentials? In most organisations the honest answer is that an engineer created a token, pasted it into an agent config, and shipped it. No review, no owner of record, no expiry. If you asked today for a list of agents in your environment with write access to production, you would get a guess.
What can it do without a human in the loop? This is a design decision that almost nobody writes down. Reading logs is different from restarting a service, which is different from deleting a resource or opening a pull request that another agent approves. If the boundary is not written, the boundary is whatever the tools happen to allow.
Who is accountable when it acts? Not who gets blamed. Who is the named person whose job includes noticing. An agent with no owner is an unowned production service that also writes.
A capability model you can actually run
You do not need a policy engine to start. You need tiers and a named owner per agent. This is the shape I would use.
Tier | What the agent can touch | Approval | Audit |
|---|---|---|---|
Read | Logs, metrics, code, tickets | Team lead | Sampled |
Propose | Opens PRs, drafts changes, no merge | Team lead | Every action logged |
Act, reversible | Restarts, scale up and down, cache flush | Engineering manager, named owner | Every action logged, owner reviews weekly |
Act, irreversible | Deletes, IAM changes, data movement, spend | No agent does this alone | Human approval per action |
The value is not the table. It is that every agent in your environment has to sit in a row, which forces someone to say out loud what a given agent is allowed to do. Most of the agents you have today were never assigned to a row, and about half of them would not survive being assigned honestly.
What to fix before you buy anything
Five things, in order, none of which need new hardware.
Inventory the credentials, not the agents. Agents are easy to hide, tokens are not. Search your secret store and your cloud IAM for non-human identities created in the last year, and map each one to a person. Anything that does not map is your real problem.
Give each agent its own identity. Agents sharing a team service account is the equivalent of everyone sharing a root password, and it removes any possibility of telling which one acted. Scope per agent, per environment, shortest lifetime you can operate with.
Log to somewhere the agent cannot write. If the agent can reach the audit trail, the audit trail is a suggestion. This is the one principle from the NVIDIA design that translates directly, and it costs nothing but a separate sink and a stricter IAM policy.
Define and test the kill switch. Every agent needs an answer to "how do we stop it in thirty seconds", and the answer has to be tested. Revoking a token is the usual mechanism, which means you need to know which token, which is why the inventory comes first.
Decide the irreversible list. Write down the actions no agent performs without a human, and make it short enough that people remember it. Deletes, IAM, anything that moves money or customer data.
That is a week of work for most teams, and it is the difference between having agents and having agents you can describe.
The honest caveats
Two things are worth saying plainly, since both cut against the announcement.
Hardware enforcement is not available to most of you, and will not be for a while. Reference designs and DPUs are a data centre purchase. If your infrastructure is somebody else's cloud, you will get this when your provider offers it, on their timeline, not yours.
And containment is not the same thing as correctness. A quarantined agent is an agent that has already done something. The controls above reduce blast radius, which matters, but none of them make an agent's judgement better. The design question underneath all of this is still which decisions you are willing to let a system make unattended, and that one has no vendor answer.
What to take to your next planning meeting
One number: how many non-human identities in your production environment have write access, and how many of those have a named owner.
If you cannot produce that number this week, it is the most useful thing your platform team could build before the end of the quarter. Everything else in agent governance, including anything you eventually buy, sits on top of knowing what is already running with your credentials.
More of our work on how infrastructure teams make these decisions sits in leadership, and the AI infrastructure guide covers the serving and inference layer these agents run on.







