Agent runtime security | 2 October 2026 | Agent engineering

A system prompt is not a security boundary.

Maddipalli Gopalakrishna · AI / ML Engineer

30 seconds · music only

Your AI agent can query databases, call APIs, run code, read files and operate tools. What actually stops it from doing something it shouldn't?

For a lot of agents today, the honest answer is a line in the system prompt.

SYSTEM: "You may read production data, but never modify it."

Guidance is not enforcement

That instruction is useful. It shapes what the model tries to do, and most of the time it works. But it's a request, not a control. Prompt injection, a misread tool result or a plain reasoning error can all lead an agent to ignore it, and nothing in the prompt can stop the call once it's made.

We would never give a normal application database-admin credentials and rely on a README that says “please don't delete anything”. An agent deserves the same treatment, arguably stricter, because its next action is decided at runtime rather than written in advance.

A system prompt tells an agent what it should do. Architecture determines what it is actually allowed to do.

How the architecture changed

Early LLM apps

  1. User
  2. Model
  3. Answer

Early agents

  1. User
  2. Agent
  3. Tools
  4. Action

Production agents

  1. User
  2. Authentication
  3. Agent
  4. Agent identity
  5. Policy engine
  6. Runtime sandbox
  7. Authorized tools / MCP
  8. Infrastructure

The production version adds audit and observability across every step, so each decision can be traced afterwards.

The layers that enforce limits

Agent identity
Every call is attributable to a specific agent, not just to the user who started it.
Scoped credentials
The agent holds the narrowest token that does the job, and it expires.
Tool-level authorization
Permission is checked per tool and per operation, with reads and writes kept separate.
Policy engine
Allow, deny or escalate is decided by code outside the model, which the model can't talk its way past.
Runtime sandbox
Code, files and network egress run inside a boundary the agent can't widen.
MCP authorization
Each MCP server checks who is calling and what they may do, rather than trusting the client.
Human approval
Consequential writes wait for a person, and the agent can propose but not commit.
Audit and evals
Every action is logged so it can be replayed, and security evals run before each release.

None of these depend on the model behaving well. That's the point: they still hold when it doesn't.

Policy decides, one action at a time

The same agent, with the same prompt, gets three different answers depending on what it is trying to do:

Requested actionPolicy result
READ DATABASE✓ Allow
DELETE DATABASE✕ Deny
UPDATE VERIFIED DATA⚠ Human approval

The deny happens before the request reaches infrastructure. The approval case is the one teams most often skip: an agent that can propose a write but not commit it is far more useful than one that can do neither, and far safer than one that can do both.

The industry is moving here

On 28 September 2026 NVIDIA announced its Open Agent Safety Platform. It's built around OpenShell, an open-source runtime that runs agents in sandboxes governed by declarative policy. It's one example of a broader shift rather than the only way to do this. Most of the controls above can be built today with ordinary identity, policy and container tooling.

If 2024 was mostly about prompt engineering and 2025 about agent engineering, the emphasis now is moving to agent infrastructure engineering: the layers around the model that decide what it can touch. These aren't clean eras, just where the hard problems have been.

Bounded consequences

The safest production agent isn't necessarily the one that follows every instruction perfectly. It's the one running inside an architecture where its mistakes have bounded consequences.

So don't just ask “How intelligent is my agent?” Also ask “How much authority should it have?”

As agents become more autonomous, should authorization and sandboxing become standard layers in every production AI architecture?