The architecture
A code interpreter is a model with one tool: "run this Python". The model writes code, a sandbox runs it, and the output goes back to the model until it can answer. Three pieces:
- The model, served through an OpenAI-compatible API with tool calling. Here: an open model on your own GPUs through gpuOS, at
https://gpuos.si/v1. - The sandbox, where the code runs: a cpuos microVM from the
pythontemplate, which ships CPython 3.13, uv, numpy, pandas and matplotlib. - The loop, about 40 lines of Python that pass tool calls from one to the other.
Nothing here is specific to gpuOS: any OpenAI-compatible endpoint works. Using an open model on your own hardware means the data the user uploads and the code that analyzes it never leave infrastructure you control.
Pick a model that calls tools well
Tool calling is the skill that matters. Models trained for it, such as Qwen3 32B, GLM-4 32B and gpt-oss 20B, produce valid tool calls and recover from errors in their own code. Small 8B models work for simple tasks but give up sooner.
On gpuOS, check the model catalog for licenses and how much VRAM each model needs. A 24 GB card runs Qwen3 32B at 4-bit. The self-hosting guide covers connecting a GPU, deploying the model and creating an API key.
Step 1: define the tool
One function, one argument. Write the code to a file instead of passing it on the command line, so quoting and multi-line code are never a problem. Return the exit code and a truncated stdout and stderr: the model needs the error message, not a 10 MB log.
import jsonimport osfrom cpuos import Sandboxfrom openai import OpenAIllm = OpenAI(base_url="https://gpuos.si/v1", api_key=os.environ["GPUOS_API_KEY"])sbx = Sandbox.create(template="python", timeout="30m")TOOLS = [{ "type": "function", "function": { "name": "run_python", "description": "Run a Python 3 script in an isolated sandbox. " "Files are in /work. Returns exit_code, stdout and stderr.", "parameters": { "type": "object", "properties": {"code": {"type": "string"}}, "required": ["code"], }, },}]def run_python(code: str) -> str: sbx.files.write("/work/main.py", code) run = sbx.exec("cd /work && python main.py", timeout="2m") return json.dumps({ "exit_code": run.exit_code, "stdout": run.stdout[-4000:], "stderr": run.stderr[-2000:], })Step 2: the loop
Send the conversation and the tool list. If the model answers with tool calls, run each one and append the result as a tool message. Repeat until the model answers in plain text, with a hard cap on the number of steps.
MAX_STEPS = 8def ask(question: str) -> str: messages = [ {"role": "system", "content": "You are a data analyst. Use run_python " "to compute every number. Save charts as PNG files in /work."}, {"role": "user", "content": question}, ] for _ in range(MAX_STEPS): reply = llm.chat.completions.create( model="qwen3-32b", messages=messages, tools=TOOLS ) msg = reply.choices[0].message messages.append(msg) if not msg.tool_calls: return msg.content for call in msg.tool_calls: args = json.loads(call.function.arguments) messages.append({ "role": "tool", "tool_call_id": call.id, "content": run_python(args["code"]), }) return "Stopped after too many steps."Step 3: files in, charts out
Upload the user's file before the first question and download whatever the model saved. Tell the model where the files are in the system prompt or the question.
with open("sales.csv", "rb") as f: sbx.files.write("/work/sales.csv", f.read())print(ask("Load /work/sales.csv, plot revenue per month to /work/chart.png " "and tell me which month grew the most."))with open("chart.png", "wb") as f: f.write(sbx.files.read("/work/chart.png"))session = sbx.pause() # resume with Sandbox.resume(session) on the next questionBecause pause keeps memory and disk, the next question can reuse files and installed packages from the previous one. Store the session id with the conversation and resume it when the user writes again.
Step 4: keep it safe and cheap
- One sandbox per conversation, never shared between users.
- Timeouts at two levels: per command (
timeout="2m") and per sandbox (timeout="30m"). - A step cap in the loop, so a model stuck on an error cannot spin forever.
- Truncated output, as in
run_python, so a runaway print does not fill the context window. - Egress: if the analysis does not need the internet, restrict it. Package registries are a common allowlist.
- Pause between turns: the user may take minutes to reply. A paused sandbox costs storage only.
Cost check: a 2 vCPU / 4 GB sandbox costs $0.11 per running hour on cpuos. A conversation that runs code for 3 minutes in total costs well under a cent, and the Free plan includes $5 of usage per month.
cpuos is in early access. The SDK calls on this page show the API shape early-access teams build against; names can still change before general availability.
Without writing the loop
Agent frameworks run this loop for you: see the OpenAI Agents SDK, LangGraph and Vercel AI SDK pages. The planned pairing between the two products goes one step further: a gpuOS workspace stores a cpuos API key, and the gateway runs code-interpreter tool calls in a cpuos sandbox, so any open model you serve gets the tool with no framework at all.