When to use this profile
Turn uploaded documents into structured text that an agent can cite. Preserve page numbers, document IDs and extraction errors instead of flattening an entire archive into one string. Text PDFs and scanned images need different pipelines: OCR requires additional binaries, languages and CPU time that should be measured independently.
Prepare a repeatable environment
- Pin parsers and include OCR software only if scanned documents are in scope.
- Validate allowed file types, total bytes, page count and archive expansion size.
- Write extracted data under /work/output and keep the source document unchanged.
python /work/extract_documents.pyReturn results the agent can use
- Text or JSON grouped by source document and page.
- Extraction warnings for encrypted, damaged or unsupported files.
- A manifest that lets the agent cite the source of each result.
Use one text PDF, one scanned page and one malformed input. Assert page references survive extraction and unsupported files return a useful error.
Resources and boundaries
Start with 2 vCPU and 4 GB of RAM, then measure peak memory and task duration on a representative fixture. These are workload planning values, not a benchmark or a provisioned configuration. Use the sandbox cost calculator to estimate running time and retained snapshots.
- Documents can contain malformed input, embedded scripts and prompt injection in visible or hidden text.
- Apply page and pixel limits before expensive rasterization or OCR.
- Remove document contents from diagnostic logs when those logs leave the user's workspace.
Keep model inference separate from this execution profile. A hosted model or gpuOS can decide the next action while the CPU environment runs it. The quickstart describes the account workflow and the proposed runtime contract.