Agent Sandboxing: Containers, gVisor and microVMs
AI Agents

Agent Sandboxing: Containers, gVisor and microVMs

An agent that runs code needs a blast radius, not trust. How containers, gVisor and microVMs differ, and the network and credential controls that matter.

Once an agent can execute code, the security question stops being about prompts and starts being about isolation. The model will eventually run something wrong — because it misread a path, because a dependency was malicious, or because injected text in a file told it to. Your job is to make that survivable.

The good news is that this is a solved problem in the general case. Running untrusted code is what public clouds do all day. The bad news is that the default developer setup — a Docker container with the repository mounted and the real environment variables present — is not the solved version.

Pick an isolation boundary deliberately

There are three tiers in common use, and the difference between them is where the kernel boundary sits.

Ordinary containers. Namespaces and cgroups over a shared host kernel. Excellent for resource limiting and process separation, and appropriate when the code inside is trusted-but-clumsy. The whole isolation story rests on the host kernel not having an exploitable bug, because every syscall from inside reaches it directly.

gVisor. A user-space kernel that intercepts syscalls and services most of them itself, so the guest talks to gVisor rather than to the host kernel. That shrinks the attack surface substantially without running a full VM. The trade is performance: broadly minimal overhead on compute-bound work, but meaningfully worse on I/O-heavy workloads, in the region of ten to thirty per cent.

MicroVMs. Firecracker and Kata give each workload its own kernel behind a hardware virtualisation boundary. This is the strongest of the three and, contrary to intuition, not slow to start: Firecracker microVMs boot in around 125 milliseconds with under 5 MiB of memory overhead per VM, and a single host can launch on the order of 150 per second. Snapshot-and-restore brings effective cold starts down further.

The rough rule: containers for code you wrote, gVisor when you want cheap hardening of mostly-trusted workloads, microVMs when the code is genuinely untrusted or when tenants must not be able to reach each other.

Isolation is not the same as sandboxing

A perfectly isolated VM with your production database credentials in its environment is not a sandbox. It is a compromised production client with good process hygiene.

The controls that determine actual blast radius are mostly about what the sandbox can reach:

  • Egress allowlist. Default deny outbound. Permit the package registry and whatever specific hosts the task needs. Exfiltration requires a channel out; removing the channel defeats the payoff even when the attack succeeds.
  • No ambient credentials. Nothing in the environment, nothing on disk, no metadata service. Inject short-lived, narrowly scoped tokens per task, and let them expire.
  • Read-only by default. Mount the workspace writable, everything else read-only. Most agent work needs to write to one directory.
  • Resource ceilings. CPU, memory, disk and process count. An unbounded fork loop is a denial of service you will pay for.
  • Ephemerality. Destroy the sandbox after the task. Reuse turns a one-off compromise into persistence.
  • Cloud metadata endpoint blocked. On a cloud host this is the classic path from "runs code" to "has instance credentials". Block it explicitly.

The network is the control that matters most

If you only implement one thing from this article, implement egress filtering.

Almost every serious agent incident needs data to leave. Whether the trigger was injected instructions, a poisoned dependency or a plain mistake, the damage step is an outbound request. An allowlist that permits your package registry and refuses everything else eliminates a whole category of outcome without materially inconveniencing legitimate work.

Package installation is the awkward exception, since registries host arbitrary code and install scripts run with the sandbox privileges. Prefer a vetted internal mirror, or install dependencies during image build and run the agent with networking off entirely.

Do not overlook rendered output

A subtle exfiltration route that bypasses your network rules: the agent produces content that your UI fetches. A markdown image whose URL embeds data will be retrieved by the viewer browser, outside the sandbox, using the viewer network.

Sandboxing the execution environment does not cover this. Sanitise agent output before rendering it, and constrain which hosts rendered content can contact.

Handling the results the agent brings back

Whatever the sandbox produces is untrusted data, and it flows straight back into the model context. Two implications:

First, tool output can carry injected instructions. Command output, file contents, and HTTP responses are all attacker-influencable in the general case. The sandbox limits what the code could do; it does not stop the text from steering the next model decision.

Second, results that leave the sandbox need review at the boundary. An agent that produces a diff should hand you a diff, not push a branch. Keep the write-back step outside the sandbox, gated by whatever check you would apply to a contribution from a stranger.

A pragmatic setup

  1. Container per task for internal, trusted-code workloads; gVisor or a microVM the moment untrusted content or code is involved.
  2. Default-deny egress with a small allowlist.
  3. No long-lived credentials inside; short-lived scoped tokens only.
  4. Writable workspace, read-only everything else, hard resource limits.
  5. Destroy after use; never reuse across tasks or tenants.
  6. Treat all output as untrusted, both as model input and as rendered content.
  7. Log every command with its arguments outside the sandbox, so the record survives the sandbox.

The measure of a good sandbox is not that nothing bad happens inside it. It is that you can let something bad happen inside it, destroy the box, and carry on.

Common questions

Is a Docker container enough to sandbox an AI agent?

For code you wrote and dependencies you vetted, usually yes. For genuinely untrusted code or multi-tenant work it is not, because containers share the host kernel and every syscall reaches it. gVisor or a microVM gives a real kernel boundary.

Are microVMs too slow to start for per-task sandboxes?

No. Firecracker microVMs boot in roughly 125 milliseconds with under 5 MiB of memory overhead each, and a host can launch on the order of 150 per second. Snapshot restore reduces effective cold start further, so per-task VMs are practical.

What is the single most valuable sandbox control?

Default-deny outbound networking. Nearly every damaging outcome requires data to leave, so an egress allowlist blocks the payoff even when the code inside is doing exactly what an attacker wanted.

Similar articles

Agent File System Access: Designing the Blast Radius
AI Agents
AI Agents·8 min read

Agent File System Access: Designing the Blast Radius

The filesystem is an agent's most consequential tool surface. Path containment, read and write asymmetry, and why files are an injection vector.

Read
Agent Network Access Control: Egress Is the Real Risk
AI Agents
AI Agents·9 min read

Agent Network Access Control: Egress Is the Real Risk

An agent with a shell and open network egress can exfiltrate anything it can read. How to scope outbound access without breaking the toolchain.

Read
Agent Shell Access Safety: Giving a Model a Terminal
AI Agents
AI Agents·9 min read

Agent Shell Access Safety: Giving a Model a Terminal

A shell tool is the most useful and most dangerous thing you can give an agent. How to grant it without handing over the whole machine.

Read