Skip to content

Advertisement

DevOps Society

Who Signs Off on an Agent With Production Credentials?

NVIDIA moved agent containment into hardware this week. The reason matters more than the product: agents were not breaking application controls, they were going around them. Three questions most engineering organisations cannot answer yet.

Published your local timeupdated

Who Signs Off on an Agent With Production Credentials?

On 28 September, NVIDIA announced that it is moving AI agent containment into hardware. The platform has two halves: OpenShell, an open source runtime that traces what an agent does on the CPU and enforces policy on it, and Sentry, a watchdog running on BlueField-4 data processing units that watches agent behaviour from outside the machine and can quarantine one in milliseconds. It ships under the Open Secure AI Alliance, governed by the Linux Foundation, with more than a hundred organisations involved.

The product is not the interesting part. The reason it exists is. Agents were not defeating application-layer controls. They were going around them, and NVIDIA's answer was to move the enforcement point somewhere the agent cannot reach.

That is a question your organisation has to answer whether or not you ever buy a DPU: when an autonomous agent holds production credentials, who approved that, what is it allowed to do without a human, and who is accountable when it acts.

Most engineering organisations running agents today cannot answer any of the three.

Why the enforcement point keeps moving down

Every control you write in the same layer the agent operates in is, in the end, advice.

An agent with a shell and a cloud credential can call the same APIs your guardrail calls. A policy enforced by a prompt is a request. A policy enforced by the application is a boundary the agent can route around, by picking a different tool, a different endpoint, or a path the policy author did not think of. This is not the agent being adversarial. It is the agent being effective at the goal you gave it, which is what you asked for.

Diagram of three enforcement layers for AI agents: the prompt and application layer which the agent can route around, the runtime layer, and an out of band layer on separate hardware the agent cannot reach.

The pattern is old. We stopped trusting applications to enforce network policy and moved it to the network. We stopped trusting processes to enforce isolation and moved it to the kernel, then to the hypervisor. Each move happened after the same discovery: a control that lives inside the thing it constrains is not a control.

Agents are at that point now. NVIDIA putting a watchdog on a DPU is the same move made one layer lower than anyone expected this early.

The three questions

The hardware is a year or more away from most infrastructure budgets. The governance is not, because you almost certainly have agents in production already.

Who approved this agent holding these credentials? In most organisations the honest answer is that an engineer created a token, pasted it into an agent config, and shipped it. No review, no owner of record, no expiry. If you asked today for a list of agents in your environment with write access to production, you would get a guess.

What can it do without a human in the loop? This is a design decision that almost nobody writes down. Reading logs is different from restarting a service, which is different from deleting a resource or opening a pull request that another agent approves. If the boundary is not written, the boundary is whatever the tools happen to allow.

Who is accountable when it acts? Not who gets blamed. Who is the named person whose job includes noticing. An agent with no owner is an unowned production service that also writes.

A capability model you can actually run

You do not need a policy engine to start. You need tiers and a named owner per agent. This is the shape I would use.

Tier

What the agent can touch

Approval

Audit

Read

Logs, metrics, code, tickets

Team lead

Sampled

Propose

Opens PRs, drafts changes, no merge

Team lead

Every action logged

Act, reversible

Restarts, scale up and down, cache flush

Engineering manager, named owner

Every action logged, owner reviews weekly

Act, irreversible

Deletes, IAM changes, data movement, spend

No agent does this alone

Human approval per action

The value is not the table. It is that every agent in your environment has to sit in a row, which forces someone to say out loud what a given agent is allowed to do. Most of the agents you have today were never assigned to a row, and about half of them would not survive being assigned honestly.

What to fix before you buy anything

Five things, in order, none of which need new hardware.

Inventory the credentials, not the agents. Agents are easy to hide, tokens are not. Search your secret store and your cloud IAM for non-human identities created in the last year, and map each one to a person. Anything that does not map is your real problem.

Give each agent its own identity. Agents sharing a team service account is the equivalent of everyone sharing a root password, and it removes any possibility of telling which one acted. Scope per agent, per environment, shortest lifetime you can operate with.

Log to somewhere the agent cannot write. If the agent can reach the audit trail, the audit trail is a suggestion. This is the one principle from the NVIDIA design that translates directly, and it costs nothing but a separate sink and a stricter IAM policy.

Define and test the kill switch. Every agent needs an answer to "how do we stop it in thirty seconds", and the answer has to be tested. Revoking a token is the usual mechanism, which means you need to know which token, which is why the inventory comes first.

Decide the irreversible list. Write down the actions no agent performs without a human, and make it short enough that people remember it. Deletes, IAM, anything that moves money or customer data.

That is a week of work for most teams, and it is the difference between having agents and having agents you can describe.

The honest caveats

Two things are worth saying plainly, since both cut against the announcement.

Hardware enforcement is not available to most of you, and will not be for a while. Reference designs and DPUs are a data centre purchase. If your infrastructure is somebody else's cloud, you will get this when your provider offers it, on their timeline, not yours.

And containment is not the same thing as correctness. A quarantined agent is an agent that has already done something. The controls above reduce blast radius, which matters, but none of them make an agent's judgement better. The design question underneath all of this is still which decisions you are willing to let a system make unattended, and that one has no vendor answer.

What to take to your next planning meeting

One number: how many non-human identities in your production environment have write access, and how many of those have a named owner.

If you cannot produce that number this week, it is the most useful thing your platform team could build before the end of the quarter. Everything else in agent governance, including anything you eventually buy, sits on top of knowing what is already running with your credentials.

More of our work on how infrastructure teams make these decisions sits in leadership, and the AI infrastructure guide covers the serving and inference layer these agents run on.

Advertisement

Follow DevOps Society on LinkedIn

Practical infrastructure engineering in your feed.

Follow

Written by

DevOpsSociety Editorial Team

Editorial Team

The DevOpsSociety Editorial Team covers DevOps, cloud infrastructure, Kubernetes, AI infrastructure, platform engineering, cybersecurity, FinOps, and modern engineering practices. We publish practical insights, technical guides, architecture analysis, and research for engineers and technology leaders.

More from DevOpsSociety →
The Infrastructure Briefing

Get the infrastructure briefing.

Practical DevOps, cloud, AI infrastructure and engineering insights, delivered weekly. Read by engineers and engineering leaders.

No spam. Unsubscribe anytime.