cpuos

Tutorials · 4 min read · updated Oct 6, 2026

Build a code interpreter with open models

Build a ChatGPT-style code interpreter with an open model on gpuOS and a cpuos sandbox: tool schema, agent loop, files, charts and limits, in Python.

The architecture

A code interpreter is a model with one tool: "run this Python". The model writes code, a sandbox runs it, and the output goes back to the model until it can answer. Three pieces:

  1. The model, served through an OpenAI-compatible API with tool calling. Here: an open model on your own GPUs through gpuOS, at https://gpuos.si/v1.
  2. The sandbox, where the code runs: a cpuos microVM from the python template, which ships CPython 3.13, uv, numpy, pandas and matplotlib.
  3. The loop, about 40 lines of Python that pass tool calls from one to the other.

Nothing here is specific to gpuOS: any OpenAI-compatible endpoint works. Using an open model on your own hardware means the data the user uploads and the code that analyzes it never leave infrastructure you control.

Pick a model that calls tools well

Tool calling is the skill that matters. Models trained for it, such as Qwen3 32B, GLM-4 32B and gpt-oss 20B, produce valid tool calls and recover from errors in their own code. Small 8B models work for simple tasks but give up sooner.

On gpuOS, check the model catalog for licenses and how much VRAM each model needs. A 24 GB card runs Qwen3 32B at 4-bit. The self-hosting guide covers connecting a GPU, deploying the model and creating an API key.

Step 1: define the tool

One function, one argument. Write the code to a file instead of passing it on the command line, so quoting and multi-line code are never a problem. Return the exit code and a truncated stdout and stderr: the model needs the error message, not a 10 MB log.

interpreter.py (part 1)
import jsonimport osfrom cpuos import Sandboxfrom openai import OpenAIllm = OpenAI(base_url="https://gpuos.si/v1", api_key=os.environ["GPUOS_API_KEY"])sbx = Sandbox.create(template="python", timeout="30m")TOOLS = [{    "type": "function",    "function": {        "name": "run_python",        "description": "Run a Python 3 script in an isolated sandbox. "        "Files are in /work. Returns exit_code, stdout and stderr.",        "parameters": {            "type": "object",            "properties": {"code": {"type": "string"}},            "required": ["code"],        },    },}]def run_python(code: str) -> str:    sbx.files.write("/work/main.py", code)    run = sbx.exec("cd /work && python main.py", timeout="2m")    return json.dumps({        "exit_code": run.exit_code,        "stdout": run.stdout[-4000:],        "stderr": run.stderr[-2000:],    })

Step 2: the loop

Send the conversation and the tool list. If the model answers with tool calls, run each one and append the result as a tool message. Repeat until the model answers in plain text, with a hard cap on the number of steps.

interpreter.py (part 2)
MAX_STEPS = 8def ask(question: str) -> str:    messages = [        {"role": "system", "content": "You are a data analyst. Use run_python "         "to compute every number. Save charts as PNG files in /work."},        {"role": "user", "content": question},    ]    for _ in range(MAX_STEPS):        reply = llm.chat.completions.create(            model="qwen3-32b", messages=messages, tools=TOOLS        )        msg = reply.choices[0].message        messages.append(msg)        if not msg.tool_calls:            return msg.content        for call in msg.tool_calls:            args = json.loads(call.function.arguments)            messages.append({                "role": "tool",                "tool_call_id": call.id,                "content": run_python(args["code"]),            })    return "Stopped after too many steps."

Step 3: files in, charts out

Upload the user's file before the first question and download whatever the model saved. Tell the model where the files are in the system prompt or the question.

interpreter.py (part 3)
with open("sales.csv", "rb") as f:    sbx.files.write("/work/sales.csv", f.read())print(ask("Load /work/sales.csv, plot revenue per month to /work/chart.png "          "and tell me which month grew the most."))with open("chart.png", "wb") as f:    f.write(sbx.files.read("/work/chart.png"))session = sbx.pause()  # resume with Sandbox.resume(session) on the next question

Because pause keeps memory and disk, the next question can reuse files and installed packages from the previous one. Store the session id with the conversation and resume it when the user writes again.

Step 4: keep it safe and cheap

  • One sandbox per conversation, never shared between users.
  • Timeouts at two levels: per command (timeout="2m") and per sandbox (timeout="30m").
  • A step cap in the loop, so a model stuck on an error cannot spin forever.
  • Truncated output, as in run_python, so a runaway print does not fill the context window.
  • Egress: if the analysis does not need the internet, restrict it. Package registries are a common allowlist.
  • Pause between turns: the user may take minutes to reply. A paused sandbox costs storage only.

Cost check: a 2 vCPU / 4 GB sandbox costs $0.11 per running hour on cpuos. A conversation that runs code for 3 minutes in total costs well under a cent, and the Free plan includes $5 of usage per month.

cpuos is in early access. The SDK calls on this page show the API shape early-access teams build against; names can still change before general availability.

Without writing the loop

Agent frameworks run this loop for you: see the OpenAI Agents SDK, LangGraph and Vercel AI SDK pages. The planned pairing between the two products goes one step further: a gpuOS workspace stores a cpuos API key, and the gateway runs code-interpreter tool calls in a cpuos sandbox, so any open model you serve gets the tool with no framework at all.

Questions

Which open model is best for a code interpreter?
Pick a model trained for tool calling with enough size to debug its own code. Qwen3 32B and GLM-4 32B fit a 24 GB GPU at 4-bit; gpt-oss 20B fits in about 14 GB.
Do I need gpuOS for this?
No. Any OpenAI-compatible endpoint works, including hosted APIs. gpuOS is the option when the model must run on your own GPUs.
Is state kept between tool calls?
Files and installed packages are, because the sandbox lives for the whole conversation. Python variables are not, since each call runs a new process. For variables that persist, run a Jupyter kernel in the sandbox.
How do I show charts to the user?
Have the model save images to a known folder, download them with files.read and display them in your UI.

Related

Give your agents a sandbox

cpuos is in early access: a Firecracker microVM per task, hosted in the EU or on your servers, billed per second and free while paused.

gpuOS · where models think

Need the model too? Run it on gpuOS

gpuOS serves open models on your own GPUs behind one OpenAI-compatible API. The model reasons on gpuOS, the agent acts in a cpuOS sandbox.