cpuos

Getting started · 4 min read · updated Oct 6, 2026

What is an AI agent sandbox?

An AI agent sandbox is an isolated, disposable computer where an agent runs the code it writes. What it must block, what agents need from it, how to pick one.

An agent is a loop, and half of it runs code

An AI agent is a loop. The model reads the task and proposes an action: run this script, open this page, edit this file. Something executes the action. The result goes back into the model's context, and the loop starts again until the task is done.

The model half of the loop needs GPUs. You can call a hosted API, or serve open models on your own hardware with gpuOS. The action half needs a computer. An AI agent sandbox is that computer: an isolated, disposable Linux machine where the agent can run commands, write files, install packages and start servers without touching anything you care about.

The short version of the brand story behind cpuos: models think on GPUs, agents act on CPUs. A sandbox is where the acting happens.

What a sandbox has to stop

Code written by a model is untrusted input. The model can make mistakes, and it can be steered by prompt injection hidden in a web page, a PDF or a GitHub issue it was asked to read. A sandbox assumes the worst and contains it.

  • Damage to the host: deleting files, filling the disk, killing processes.
  • Secret theft: reading environment variables, SSH keys, cloud credentials, or the cloud metadata endpoint at 169.254.169.254.
  • Lateral movement: scanning and calling services on your private network, such as a database on 10.0.0.0/8.
  • Abuse from your IP addresses: crypto mining, spam over port 25, attacks on third parties.
  • Resource exhaustion: fork bombs, memory leaks, infinite loops that block every other user.
  • Escape: breaking out of the isolation boundary to reach the host or other tenants.

What an agent needs from it

Blocking everything is easy. A useful sandbox also gives the agent a real machine to work with:

  • Fast start: agents create sandboxes on demand, often one per task. A start time measured in seconds adds up across a session.
  • Exec with streaming: run a command, stream stdout and stderr, get the exit code, enforce a timeout.
  • Files: upload a dataset, download the chart or the patch the agent produced.
  • Ports: expose a dev server or a notebook so a user can see a live preview.
  • A browser: many agents browse, fill forms or test web apps.
  • Pause and resume: agents spend most of their wall-clock time waiting for the model or the user. Paying for an idle VM is waste.
  • Limits and an audit trail: CPU, memory, disk and session length per sandbox, and a log of every command and outbound connection.

Levels of isolation

ApproachShares the host kernelTypical startFit for untrusted code
exec() or a subprocess in your appYes, and your processInstantNo
Docker or another OCI containerYesSub-second to secondsOnly with heavy hardening, single tenant
gVisor (user-space kernel)Partly: a small syscall surfaceSub-second to secondsGood, with some compatibility gaps
Firecracker microVMNo: own guest kernel on KVMHundreds of ms from a snapshotYes, multi-tenant
Full VM (QEMU, cloud instance)NoTens of secondsYes, but slow and heavy

The guide on running LLM-generated code safely goes through each level in detail, and Firecracker vs Docker compares the two most common choices.

How cpuos does it

Each cpuos sandbox is a Firecracker microVM with its own Linux kernel, started through Firecracker's jailer with seccomp filters and cgroups. Templates (Python, Node, Browser, Dev box or your own image) are booted once and snapshotted, so a new sandbox restores from that snapshot instead of booting.

TypeScript
import { Sandbox } from "@cpuos/sdk"const sbx = await Sandbox.create({ template: "python", timeout: "15m" })const run = await sbx.exec("python analysis.py", {  onStdout: (line) => console.log(line),  timeout: "2m",})const chart = await sbx.files.read("/work/chart.png")const id = await sbx.pause() // CPU and RAM billing stop here

Outbound traffic follows a policy: internet on, an allowlist, or off. Private ranges, SMTP and the metadata address are always blocked. Usage is billed per second while a sandbox runs, and a paused sandbox only pays for snapshot storage.

cpuos is in early access. The SDK calls on this page show the API shape early-access teams build against; names can still change before general availability.

When you do not need one

If your agent only calls functions you wrote, such as search_orders(customer_id) or create_ticket(title), and never runs code or shell commands the model generated, you do not need a sandbox. You need input validation and authorization on those functions.

You need a sandbox as soon as the model writes code that runs, installs packages, browses arbitrary sites or edits a repository. That includes the popular "code interpreter" tool: see how to build one with open models.

Questions

Is an AI agent sandbox the same as a code interpreter?
A code interpreter is one use of a sandbox: the model writes Python, the sandbox runs it and returns the output. Agents also use sandboxes to browse, run test suites, start dev servers and edit repositories.
Can I use a Docker container as a sandbox?
For your own trusted code, yes. For code written by a model on behalf of many users, a container shares the host kernel, so one kernel or runtime bug can expose the host and every other tenant. MicroVMs remove that shared kernel.
Does the sandbox need a GPU?
Usually not. The model needs the GPU; the code the agent runs is mostly CPU work like data analysis, builds, tests and browsing. Run the model on GPUs, for example on gpuOS, and the actions in CPU sandboxes.
How long does a sandbox live?
As long as the task. On cpuos, sessions last up to 1 hour on Free, 24 hours on Team and 7 days on Business, and a paused sandbox can be resumed later.

Related

Give your agents a sandbox

cpuos is in early access: a Firecracker microVM per task, hosted in the EU or on your servers, billed per second and free while paused.

gpuOS · where models think

Need the model too? Run it on gpuOS

gpuOS serves open models on your own GPUs behind one OpenAI-compatible API. The model reasons on gpuOS, the agent acts in a cpuOS sandbox.