Start with the current CPU execution contract
A Python job needs an explicit CPU budget, memory ceiling and execution deadline. Those limits keep one approved script from consuming the worker indefinitely. They do not prove that code is trustworthy or that its output is correct. The cpuOS pilot runs trusted team code in restricted Docker containers sharing the host kernel, so allocate a worker appropriate to that trust boundary.
Scroll horizontally to see every column.
| Request field or boundary | Current contract | Planning consequence |
|---|---|---|
| template | python or node | Python 3.13 and Node 24 standard-library jobs. |
| vcpu | Integer 1–2; default 1 | Request only CPU capacity the task can use. |
| memoryMb | Integer 128–2048 MiB; default 256 | Include interpreter, input and intermediate objects. |
| timeoutSeconds | Integer 1–120; default 30 | Set a bounded script execution deadline. |
| Complete code | At most 64 KiB UTF-8 | Inline input shares the same source budget. |
| Captured output | 64 KiB per stdout/stderr stream | Return a compact aggregate and check outputTruncated. |
Use the quickstart to connect a worker, then the Python jobs client to submit real requests. Resource fields belong in the application-controlled request body. A model should not be able to raise them beyond your task policy merely by proposing larger values.
A CPU quota does not create application parallelism
A CPU limit controls how much CPU time the container can consume. It does not make a serial Python loop execute twice as fast, reserve an otherwise idle host, or promise a particular completion latency. Docker resource constraints describe CPU scheduling limits and memory ceilings at the container level.
Begin a small standard-library aggregation with one CPU. If an implementation has actual parallel work, measure that implementation before requesting two CPUs. Host load, interpreter startup, algorithm choice and worker availability all affect observed duration. Prefer reducing unnecessary allocations or repeated calculations before increasing the request.
Measure input and intermediate allocations separately
JSON source bytes are only part of memory use. Decoding creates Python objects; storing transformed copies and serializing the result can create more. A list of every intermediate value requires more space than a running aggregate. Bound row counts, string lengths and nested structures in addition to the total encoded input size.
import jsonimport timeimport tracemalloctracemalloc.start()started = time.perf_counter()# Compute an aggregate without storing a list of all squared values.count = 10_000total = sum(value * value for value in range(count))elapsed = time.perf_counter() - started_, peak = tracemalloc.get_traced_memory()tracemalloc.stop()print(json.dumps({ "count": count, "sum_squares": total, "elapsed_seconds": elapsed, "python_traced_peak_bytes": peak,}, allow_nan=False, sort_keys=True))Run python3 profile-job.py locally. The invariant result is count: 10000 and sum_squares: 333283335000; elapsed time and allocation peak vary. tracemalloc measures traced Python allocations. Its peak is not total process RSS or total container memory, so it cannot alone justify an exact memoryMb setting.
Test representative maximum inputs on your actual pilot worker with margin for interpreter and runtime overhead. If a task cannot fit within 2048 MiB, redesign or partition the workload rather than assuming a larger unsupported allocation. Keep measurement diagnostics concise and separate from the machine-readable business result.
Execution, queue and client deadlines solve different problems
Scroll horizontally to see every column.
| Deadline | What it limits | Who controls it |
|---|---|---|
| Job timeoutSeconds | The worker execution attempt | Application request, within 1–120 seconds. |
| Client waiting budget | How long the application waits across HTTP requests and queue delay | Your application or supervising process. |
| HTTP request timeout | How long one request may wait under its client's timeout semantics | Your HTTP client. |
| Queue expiry / worker lease | Stale accepted work or interrupted execution | The service's current lifecycle policy. |
A thirty-second execution timeout does not mean the whole workflow returns within thirty seconds. A job can wait in the queue before execution, and the client needs time to observe the terminal result. Set the business waiting policy deliberately. If the client stops waiting, cancel known unfinished work and preserve the ID when cancellation cannot be confirmed.
The sample uses time.perf_counter for a local elapsed measurement. Do not turn one measured duration into a promised production deadline. Follow asynchronous job handling for recovery and cancellation.
Control queue pressure before adding more jobs
An eligible online worker must report Docker availability, a recent heartbeat and enough CPU and memory for the request. cpuOS rejects submission with no_online_nodes when no such worker exists. Eligibility does not mean that a job will start immediately. Check worker readiness and the actual queued/running records rather than treating advertised host cores as instant spare capacity.
The workspace queue currently accepts up to fifty queued jobs. A full queue yields queue_full; changing API keys does not create a separate queue. Maintain a bounded submission window in your application, reduce needless polling and wait for terminal jobs before admitting more batches. Use the worker operations guide when heartbeat or Docker readiness is unclear.
Diagnose the attempt before changing its limits
- Require completed status, exitCode zero and a checked output contract before accepting success.
- Inspect stderr and the execution error for failed attempts. A nonzero exit alone does not prove memory exhaustion.
- Treat outputTruncated as an incomplete capture, then reduce the result or log volume.
- For worker_lost or queue_expired, repair availability and review the attempt before requesting another execution.
- For an expensive valid input, profile the algorithm and reduce batch size before relaxing a deadline.
Keep each retry tied to a recorded reason. Increasing memory after a syntax error or changing the timeout after an offline-worker rejection will not address the cause. Batch Python jobs explain application-side partitioning within the current source, time and output bounds.
Budget model context separately from Python execution
An agent workflow has two resource domains: model inference on GPUs and calculation on a CPU worker. A larger model context consumes inference memory, while a larger inline Python dataset consumes source bytes and CPU memory. gpuOS context window and KV cache planning covers the inference side. Return a checked summary to the model so both the job output and the next prompt remain bounded.