The browser is the riskiest tool an agent has
A browser agent visits pages nobody vetted. Every page is input to the model, so every page is a chance for prompt injection: hidden text that tells the agent to download a file, paste a token somewhere or visit another site. And Chromium itself is a large program that parses hostile content all day.
- Downloads land on whatever disk the browser runs on.
- Cookies and sessions in the browser profile are readable by anything else on that machine.
- Internal URLs such as admin panels on your private network are reachable if the browser is.
- Chromium's own sandbox is often disabled in containers with
--no-sandbox, because it needs kernel features that container runtimes restrict. That removes a layer exactly where you need it.
Running the browser in a microVM puts a kernel boundary around all of it. Inside the VM, Chromium can keep its own sandbox, and nothing on the VM matters once the task is over.
Two ways agents use a browser
- Script mode: the model writes a Playwright script for the whole task (log in, open the report, export the CSV), the sandbox runs it, and the result goes back to the model. Fast and cheap, good for structured sites and QA.
- Step mode: the model drives the browser one action at a time (click, type, scroll, screenshot), looking at the page after each step. Slower, but it handles pages the model has never seen.
The cpuos browser template ships Chromium, Playwright, a CDP endpoint and a live view, and starts in about 650 ms from its snapshot. Script mode needs only exec and files. Step mode uses the CDP endpoint; its SDK call is in the early-access docs.
Script mode with Playwright
The model writes the script, you write it to a file in the sandbox and run it. Screenshots and downloads stay in the sandbox until you read them.
import { chromium } from "playwright"const browser = await chromium.launch()const page = await browser.newPage()await page.goto("https://news.ycombinator.com")const titles = await page.locator(".titleline > a").allTextContents()await page.screenshot({ path: "/work/page.png", fullPage: true })console.log(JSON.stringify(titles.slice(0, 10)))await browser.close()import { Sandbox } from "@cpuos/sdk"const sbx = await Sandbox.create({ template: "browser", timeout: "10m" })await sbx.files.write("/work/task.mjs", script)const run = await sbx.exec("node /work/task.mjs", { timeout: "2m" })const screenshot = await sbx.files.read("/work/page.png")await sbx.pause()Return run.exitCode and the truncated output to the model. If the script failed, the error message and the screenshot are usually enough for the model to fix its selector and try again.
Step mode and screenshots
In step mode, each tool call is one browser action, and the model decides the next action from the page state. The page state is either the accessibility tree as text (cheaper, works with any model) or a screenshot (needs a vision model).
For screenshots you need a model that reads images. The gpuOS model catalog lists a Qwen3 VL 8B vision model that runs on a single GPU and is served through the same OpenAI-compatible API as text models, so the agent sends the screenshot as an image message.
Hardening checklist
- One sandbox per task, paused or discarded at the end. Never reuse a browser profile across users.
- Egress allowlist for the sites the task needs, when you know them. Private ranges and the metadata address are blocked on cpuos regardless.
- No real credentials in the sandbox when you can avoid them. Prefer test accounts or short-lived session cookies uploaded for one task.
- Timeouts per command and per sandbox, so a page that never finishes loading cannot hold resources.
- Keep the evidence: screenshots and the audit log of outbound connections show what the agent actually did.
- Pause while the model thinks: step mode spends most of its time waiting for the model. Pausing stops CPU and RAM billing.
cpuos is in early access. The SDK calls on this page show the API shape early-access teams build against; names can still change before general availability.