cpuos

Comparisons · 4 min read · updated Oct 6, 2026

Firecracker vs Docker for AI agents

Firecracker microVMs vs Docker containers for running AI agent code: kernel isolation, startup time, snapshots, density, operations and when each makes sense.

The short answer

Use Docker when the code is yours or your team's and runs on a machine you dedicate to it. Use Firecracker microVMs when the code is written by a model, for users you do not control, on hosts shared with other tenants. The difference is the kernel: containers share the host's, microVMs each get their own.

How each one isolates

Docker

A container is a set of Linux processes with their own namespaces (PID, mount, network, user), limited by cgroups. Docker's default seccomp profile blocks a few dozen dangerous syscalls and allows the rest. Everything else in the kernel, including file systems, the network stack and drivers, is shared with the host and every other container on it.

Firecracker

A microVM runs its own guest kernel on KVM. Code inside talks to the guest kernel, never to the host's. To reach the host, an attacker has to break KVM or the Firecracker process, which is written in Rust, emulates only virtio network, block and vsock devices plus a serial console, and runs under the jailer with seccomp filters, cgroups and an unprivileged user.

Side by side

Docker (runc)Firecracker microVM
KernelShared with the hostOne guest kernel per sandbox
Escape surfaceThe host kernel's syscalls, the runtimeKVM and a small virtio device model
Cold startUsually under a second with a cached imageAbout 125 ms to user space for a minimal guest
Start from a snapshotNot built in (CRIU is experimental)Built in: memory and device state restore
Memory overheadA few MB per containerUnder 5 MiB per VMM, plus the guest's own memory
ImagesOCI images and DockerfilesA kernel and a root file system, often built from an OCI image
Host requirementAny Linux host/dev/kvm: bare metal or nested virtualization
GPUYes, with the NVIDIA Container ToolkitNo GPU passthrough
ToolingHuge ecosystemLow level: you build networking, images and orchestration

The GPU row matters less than it looks. In an agent, the GPU work is the model, and it does not need to live next to the sandbox. Serve the model on GPUs, for example with gpuOS on your own cards, and keep the sandboxes on CPU hosts.

Startup and snapshots in practice

Agents create sandboxes all the time, so start time is part of the user experience. With Firecracker, the trick is to boot each template once, wait until its services are ready, and take a snapshot. A new sandbox restores that snapshot: Firecracker maps the memory file and loads pages as the guest touches them, so restore time does not grow with the template's memory size.

The same mechanism gives you pause and resume. Pausing writes the sandbox's memory and disk to a snapshot and frees the CPU and RAM. Resuming continues with processes, open files and the Python session exactly where they were. Docker can stop and start a container, but running processes do not survive that without CRIU.

Density and cost

Containers are denser: they share the page cache and the kernel, and an idle container costs almost nothing. A microVM reserves its guest memory, so RAM is the real limit on how many sandboxes fit on a host.

Two things close most of the gap for agents. CPU can be oversubscribed because agent sandboxes are idle most of the time, waiting for the model. And paused sandboxes free their RAM entirely, keeping only a snapshot on disk. On cpuos, a 2 vCPU / 4 GB sandbox costs $0.11 per running hour, and a paused one only pays snapshot storage.

Middle grounds

  • gVisor (`runsc`): a drop-in Docker runtime with a user-space kernel. Stronger than plain runc, no KVM needed, slower on syscall-heavy work.
  • Kata Containers: runs each container inside a lightweight VM (with QEMU, Cloud Hypervisor or Firecracker) behind the OCI interface, so Kubernetes and Docker workflows stay the same.
  • A managed sandbox API: someone else runs the microVMs, networking, snapshots and cleanup. You call create, exec and pause.

Keep your Dockerfiles

Choosing microVMs does not mean giving up Docker as a build tool. cpuos turns an OCI image or a Dockerfile into a microVM root file system, boots it once and snapshots it as a custom template (Team plan and up). You keep building images the way you do today; the runtime underneath is a VM.

TypeScript
import { Sandbox } from "@cpuos/sdk"// "devbox" ships git, Python, Node, Go and Rust toolchainsconst sbx = await Sandbox.create({ template: "devbox", vcpu: 4, memory: "8GB" })await sbx.exec("git clone https://github.com/acme/api /work/api")const tests = await sbx.exec("cd /work/api && make test", { timeout: "10m" })

cpuos is in early access. The SDK calls on this page show the API shape early-access teams build against; names can still change before general availability.

Questions

Is Firecracker slower than Docker?
Once running, CPU-bound code runs at close to native speed in both. Firecracker cold boots are fast, and restoring a snapshot is faster still, so start time is comparable for agent workloads.
Can Firecracker run Docker images?
Not directly. An OCI image is converted into a root file system and booted with a guest kernel. Sandbox services such as cpuos do this conversion for you.
Why not run the model and the sandbox on the same GPU machine?
Firecracker has no GPU passthrough, and mixing untrusted code with your model server widens the blast radius. Keep the model on a GPU host and the sandboxes on separate CPU hosts.

Related

Give your agents a sandbox

cpuos is in early access: a Firecracker microVM per task, hosted in the EU or on your servers, billed per second and free while paused.

gpuOS · where models think

Need the model too? Run it on gpuOS

gpuOS serves open models on your own GPUs behind one OpenAI-compatible API. The model reasons on gpuOS, the agent acts in a cpuOS sandbox.