cpuos

Recipe · python base

Document processing sandbox recipe

A Python recipe for PDF and office-document extraction, with page limits, preserved source references and separate OCR resource planning.

When to use this profile

Turn uploaded documents into structured text that an agent can cite. Preserve page numbers, document IDs and extraction errors instead of flattening an entire archive into one string. Text PDFs and scanned images need different pipelines: OCR requires additional binaries, languages and CPU time that should be measured independently.

Prepare a repeatable environment

  • Pin parsers and include OCR software only if scanned documents are in scope.
  • Validate allowed file types, total bytes, page count and archive expansion size.
  • Write extracted data under /work/output and keep the source document unchanged.
Command inside a prepared Linux guest
python /work/extract_documents.py

Return results the agent can use

  • Text or JSON grouped by source document and page.
  • Extraction warnings for encrypted, damaged or unsupported files.
  • A manifest that lets the agent cite the source of each result.

Use one text PDF, one scanned page and one malformed input. Assert page references survive extraction and unsupported files return a useful error.

Resources and boundaries

Start with 2 vCPU and 4 GB of RAM, then measure peak memory and task duration on a representative fixture. These are workload planning values, not a benchmark or a provisioned configuration. Use the sandbox cost calculator to estimate running time and retained snapshots.

  • Documents can contain malformed input, embedded scripts and prompt injection in visible or hidden text.
  • Apply page and pixel limits before expensive rasterization or OCR.
  • Remove document contents from diagnostic logs when those logs leave the user's workspace.

Keep model inference separate from this execution profile. A hosted model or gpuOS can decide the next action while the CPU environment runs it. The quickstart describes the account workflow and the proposed runtime contract.

Questions

Is this a separate built-in template?
This is a workload recipe based on the python profile. It adds preparation and execution guidance, not a separate built-in image or a new backend runtime.
Can I run the command now?
The command runs in a Linux environment where the listed dependencies and input files are prepared. cpuOS SDK examples describe a proposed contract; confirm runtime access and package versions before integration.
How should I choose CPU and memory?
Use the starting profile to run a representative fixture, measure peak memory and elapsed time, and add room for package installation, worker processes and larger inputs. Enforce a task timeout separately.

Related guides

Plan your document processing task

Create a workspace and choose a profile. Connect an execution backend before running code.

gpuOS · where models think

Need the model too? Run it on gpuOS

gpuOS serves open models on your own GPUs behind one OpenAI-compatible API. The model reasons on gpuOS, the agent acts in a cpuOS sandbox.