cpuos

Tutorials · 5 min read · updated Oct 7, 2026

CPU job resource limits: memory, timeouts and queue capacity

Choose CPU, memory and timeout limits for Python jobs. Profile bounded scripts, distinguish queue time from execution and diagnose worker capacity in cpuOS.

The current pilot runs trusted Python and Node jobs on your Docker worker. Containers share its kernel. Browser and repository workflows need capabilities beyond this pilot.

On this page

Start with the current CPU execution contract

A Python job needs an explicit CPU budget, memory ceiling and execution deadline. Those limits keep one approved script from consuming the worker indefinitely. They do not prove that code is trustworthy or that its output is correct. The cpuOS pilot runs trusted team code in restricted Docker containers sharing the host kernel, so allocate a worker appropriate to that trust boundary.

Scroll horizontally to see every column.

Request field or boundaryCurrent contractPlanning consequence
templatepython or nodePython 3.13 and Node 24 standard-library jobs.
vcpuInteger 1–2; default 1Request only CPU capacity the task can use.
memoryMbInteger 128–2048 MiB; default 256Include interpreter, input and intermediate objects.
timeoutSecondsInteger 1–120; default 30Set a bounded script execution deadline.
Complete codeAt most 64 KiB UTF-8Inline input shares the same source budget.
Captured output64 KiB per stdout/stderr streamReturn a compact aggregate and check outputTruncated.

Use the quickstart to connect a worker, then the Python jobs client to submit real requests. Resource fields belong in the application-controlled request body. A model should not be able to raise them beyond your task policy merely by proposing larger values.

A CPU quota does not create application parallelism

A CPU limit controls how much CPU time the container can consume. It does not make a serial Python loop execute twice as fast, reserve an otherwise idle host, or promise a particular completion latency. Docker resource constraints describe CPU scheduling limits and memory ceilings at the container level.

Begin a small standard-library aggregation with one CPU. If an implementation has actual parallel work, measure that implementation before requesting two CPUs. Host load, interpreter startup, algorithm choice and worker availability all affect observed duration. Prefer reducing unnecessary allocations or repeated calculations before increasing the request.

Measure input and intermediate allocations separately

JSON source bytes are only part of memory use. Decoding creates Python objects; storing transformed copies and serializing the result can create more. A list of every intermediate value requires more space than a running aggregate. Bound row counts, string lengths and nested structures in addition to the total encoded input size.

profile-job.py: local standard-library measurement fixture
import jsonimport timeimport tracemalloctracemalloc.start()started = time.perf_counter()# Compute an aggregate without storing a list of all squared values.count = 10_000total = sum(value * value for value in range(count))elapsed = time.perf_counter() - started_, peak = tracemalloc.get_traced_memory()tracemalloc.stop()print(json.dumps({    "count": count,    "sum_squares": total,    "elapsed_seconds": elapsed,    "python_traced_peak_bytes": peak,}, allow_nan=False, sort_keys=True))

Run python3 profile-job.py locally. The invariant result is count: 10000 and sum_squares: 333283335000; elapsed time and allocation peak vary. tracemalloc measures traced Python allocations. Its peak is not total process RSS or total container memory, so it cannot alone justify an exact memoryMb setting.

Test representative maximum inputs on your actual pilot worker with margin for interpreter and runtime overhead. If a task cannot fit within 2048 MiB, redesign or partition the workload rather than assuming a larger unsupported allocation. Keep measurement diagnostics concise and separate from the machine-readable business result.

Execution, queue and client deadlines solve different problems

Scroll horizontally to see every column.

DeadlineWhat it limitsWho controls it
Job timeoutSecondsThe worker execution attemptApplication request, within 1–120 seconds.
Client waiting budgetHow long the application waits across HTTP requests and queue delayYour application or supervising process.
HTTP request timeoutHow long one request may wait under its client's timeout semanticsYour HTTP client.
Queue expiry / worker leaseStale accepted work or interrupted executionThe service's current lifecycle policy.

A thirty-second execution timeout does not mean the whole workflow returns within thirty seconds. A job can wait in the queue before execution, and the client needs time to observe the terminal result. Set the business waiting policy deliberately. If the client stops waiting, cancel known unfinished work and preserve the ID when cancellation cannot be confirmed.

The sample uses time.perf_counter for a local elapsed measurement. Do not turn one measured duration into a promised production deadline. Follow asynchronous job handling for recovery and cancellation.

Control queue pressure before adding more jobs

An eligible online worker must report Docker availability, a recent heartbeat and enough CPU and memory for the request. cpuOS rejects submission with no_online_nodes when no such worker exists. Eligibility does not mean that a job will start immediately. Check worker readiness and the actual queued/running records rather than treating advertised host cores as instant spare capacity.

The workspace queue currently accepts up to fifty queued jobs. A full queue yields queue_full; changing API keys does not create a separate queue. Maintain a bounded submission window in your application, reduce needless polling and wait for terminal jobs before admitting more batches. Use the worker operations guide when heartbeat or Docker readiness is unclear.

Diagnose the attempt before changing its limits

  • Require completed status, exitCode zero and a checked output contract before accepting success.
  • Inspect stderr and the execution error for failed attempts. A nonzero exit alone does not prove memory exhaustion.
  • Treat outputTruncated as an incomplete capture, then reduce the result or log volume.
  • For worker_lost or queue_expired, repair availability and review the attempt before requesting another execution.
  • For an expensive valid input, profile the algorithm and reduce batch size before relaxing a deadline.

Keep each retry tied to a recorded reason. Increasing memory after a syntax error or changing the timeout after an offline-worker rejection will not address the cause. Batch Python jobs explain application-side partitioning within the current source, time and output bounds.

Budget model context separately from Python execution

An agent workflow has two resource domains: model inference on GPUs and calculation on a CPU worker. A larger model context consumes inference memory, while a larger inline Python dataset consumes source bytes and CPU memory. gpuOS context window and KV cache planning covers the inference side. Return a checked summary to the model so both the job output and the next prompt remain bounded.

Questions

Does requesting two CPUs double Python job speed?
No. A CPU quota is a resource limit, not automatic parallelism or a latency guarantee. Measure your implementation on the worker; a serial calculation may not benefit from the larger request.
Is timeoutSeconds the maximum time my application waits?
No. It limits the execution attempt. Queue delay, HTTP calls and polling require a separate client waiting policy, and stopping the client does not automatically cancel the accepted job.
Can tracemalloc prove a script fits the container memory limit?
No. It measures traced Python allocations, not the complete process or container footprint. Test representative maximum inputs on the worker with margin and review execution failures.

Related

Connect a worker and run a job

Start with a small trusted Python or Node task, fixed limits and an expected result.

gpuOS · where models think

Need the model too? Run it on gpuOS

gpuOS serves open models on your own GPUs behind one OpenAI-compatible API. Your application can submit authorized actions to cpuOS jobs.