cpuos

Tutorials · 9 min read · updated Oct 7, 2026

Operate a self-hosted CPU worker: readiness, queues and updates

Operate a cpuOS Linux worker with systemd. Check readiness, diagnose queued jobs, plan resource limits and update or reconnect without exposing tokens.

The current pilot runs trusted Python and Node jobs on your Docker worker. Containers share its kernel. Browser and repository workflows need capabilities beyond this pilot.

On this page

Know what you are operating

A cpuOS worker executes bounded Python and Node jobs on a machine you allocate. The hosted control plane accepts work, stores code and output, and tracks job state. This guide starts after the quickstart: it covers readiness checks, queue diagnosis, maintenance and credential replacement.

Scroll horizontally to see every column.

Managed worker requirementOperational consequence
Linux with a running systemd managerThe managed service is cpuos-worker; this procedure is for that Linux installation.
Node.js 24 or later, installed system-wideUse /usr/bin/node, /usr/local/bin/node or /bin/node. Home-directory runtime managers are not accepted by the installer.
Rootful Docker with Linux containersThe installer requires a docker-group-owned /var/run/docker.sock and checks access as the service account.
flock from util-linux and the installer prerequisitesThe worker rejects local duplicate processes for the same node token; the installer checks required commands before changing files.
One process per node token across all hostsEnroll a separate node for another worker. A local duplicate check is not a cross-host coordinator.
One running job per workerMore host CPUs do not make this worker execute multiple jobs in parallel.

The service account belongs to the Docker group. Docker daemon access grants root-level control of the host, even though submitted containers run with restrictions. Use a purpose-allocated machine and trusted team code. Do not open a remote Docker API or make the socket world-writable to repair access. Docker Engine security.

The control plane is EU-hosted; the worker runs wherever you place it. Code and results travel through that hosted service. Connecting your own worker does not self-host the control plane, create an air-gapped system or provide Firecracker isolation.

Start with read-only service and Docker checks

Run these commands on the managed Linux worker host. They report status without stopping the service, running a job or printing its credential file. A nonzero status from a failed service is evidence to inspect, not a reason to restart immediately.

Inspect the managed worker without restarting it
sudo systemctl status cpuos-worker --no-pagersudo systemctl show cpuos-worker \  --property=ActiveState --property=SubState --property=Result \  --property=ExecMainStatus --property=NRestartssudo journalctl -u cpuos-worker -n 50 --no-pagercommand -v nodenode --versioncommand -v flocksudo stat --format '%A %U %G %n' /var/run/docker.socksudo runuser -u cpuos-worker -- docker \  --host unix:///var/run/docker.sock info \  --format 'OS={{.OSType}} CPUs={{.NCPU}} MemoryBytes={{.MemTotal}}'

Expect Docker's OS to be linux and its daemon to be reachable by cpuos-worker. Compare the Node binary path with the supported system-wide locations. Your interactive shell can resolve a different Node version from the service, so a working home-directory installation is not sufficient. Node.js 24 documentation.

Check Nodes in the dashboard as well as systemd. An active service process can still be preparing images or failing to communicate with the control plane. Initial runtime-image pulls need outbound host connectivity; job containers themselves have networking disabled.

Do not print /etc/cpuos/worker.env, process environments or shell traces in a support report. The root-owned 0600 file contains the node token. Review logs before sharing: hostnames and job IDs can identify your infrastructure and work.

Separate process health, readiness and job eligibility

Scroll horizontally to see every column.

ObservationWhat it meansWhat it does not prove
Service active, runtime images preparingThe worker can report preparation heartbeats while pulling images; it does not accept a job during that phase.That Docker is ready to execute work.
Node onlineIts token is not revoked, Docker is reported available and its heartbeat is newer than 40 seconds.A reserved execution slot, available free RAM, guaranteed latency or an availability SLA.
Node offlineThe online check failed, for example due to a stale heartbeat or unavailable Docker.That code never started, or that a restart is safe for an active attempt.
Node revokedThe credential no longer authorizes worker polling.That resubmitting the same job will repair the worker connection.

Job admission also compares requested CPU and memory with the capacity reported by an online worker. That report is based on host and Docker capacity, not a continuously measured free-memory budget. A machine already under memory pressure can pass the eligibility check and still fail to execute the workload.

Use node readiness to decide whether submission is possible, then inspect job lifecycle and host measurements to decide whether this workload can run reliably. The pilot supplies neither a liveness promise nor a measured jobs-per-second commitment.

Plan resource limits and queue behavior separately

Scroll horizontally to see every column.

Current contractPlanning rule
1 or 2 CPUs; 128–2048 MiB per jobStart with 1 CPU and 256 MiB. Increase only after checking realistic inputs and host pressure.
1–120 second execution deadline; default 30 secondsThis bounds an execution attempt, not end-to-end time including queueing and image preparation.
One active job per workerBudget queued work using measured job durations and preparation time, rather than multiplying CPU count by a concurrency assumption.
50 queued jobs per workspaceA full queue rejects new submissions with queue_full; it is not an autoscaling signal that cpuOS will act on.
Queued attempts expire after 10 minutesInspect queue_expired results and decide whether another intended execution is appropriate.
60 API requests or submissions per minute per workspaceCoordinate submission and polling clients. Creating another API key does not create a separate capacity bucket.

Docker applies CPU and memory limits to job containers. The worker also disables swap for the configured container memory budget. Resource limits can constrain consumption; they do not reserve unused host RAM or make a memory-heavy transformation fit. Inspect Docker/kernel diagnostics and test the input sizes you intend to process. Docker resource constraints.

cpuOS usage totals multiply worker-reported duration by the requested CPU and memory allocations. They are allocation-based accounting values, not measurements of actual host CPU utilization or resident memory. Keep host monitoring separate when sizing a worker.

For workloads to measure, use the CSV analysis recipe and JSON validation recipe. Vary input size while checking result correctness, duration and failure behavior. A small synthetic example establishes a baseline, not production capacity.

Use the symptom to choose the next check

Scroll horizontally to see every column.

SymptomCheckNext action
Service active but node not readyRecent worker journal, preparation messages, image-pull connectivity and Docker daemon access.Resolve the observed prerequisite or connection failure; wait for image preparation rather than sending repeated jobs.
503 no_online_nodes on submissionNode heartbeat, Docker availability and reported CPU/memory against the requested limits.Restore an eligible worker or choose a smaller request only if its measured workload fits the smaller limits.
Job remains queuedOther active jobs, eligible worker capacity, heartbeat and the job's creation time.Wait within your own budget, reduce incoming submissions, or cancel the particular queued job. Do not assume a second copy will start sooner.
429 queue_full or rate_limit_exceededThe structured error code, queued count, and how often applications submit or poll.Slow the relevant clients. Cancel unnecessary queued work intentionally; extra keys do not bypass workspace limits.
Failed result with worker_lostJob ID and result, worker service journal, host or network interruption around the attempt.Establish whether effects need review before creating a new attempt. Lost jobs are not automatically replayed.
Deadline exceeded or nonzero exit codeTerminal error, stderr, outputTruncated, input size, allocated memory and host memory pressure.Correct the observed failure and re-evaluate limits. An exit code alone is not proof of an out-of-memory cause.
Worker exited with status 77Authentication/revocation message in its journal and node state in the workspace.Enroll a new node and reconnect with its token. The managed service deliberately does not restart-loop on this exit.
Duplicate-worker or configuration failureA manual worker and the managed service sharing a node token, or an unsupported prerequisite.Choose one process for that node identity after reviewing active work. Do not remove locks while another worker may still run.

For a support report, collect timestamps, the job ID, structured error code, worker version and the relevant non-secret journal excerpt. Include whether the attempt was cold or warm and whether a maintenance action occurred. Redact submitted code and outputs when they contain private data.

Drain application traffic before an update

  1. Stop new submissions in the applications and team workflow that feed this workspace. There is no drain API or maintenance switch in the current pilot.
  2. Inspect queued and running jobs in the dashboard, or use authenticated GET /v1/jobs and GET /v1/jobs/:id. Follow list pagination when present so a first page does not hide an older active attempt.
  3. Choose a maintenance wait budget. Wait for completed, failed or cancelled status and inspect the attempt's result. A running cancellation is asynchronous: verify worker acknowledgement/result or loss handling before treating the worker as idle.
  4. If you decide to cancel a particular queued or running attempt, use its actual ID with DELETE /v1/jobs/:id. Cancellation does not automatically submit a replacement.
  5. Download and inspect the current installer. Run --update only after deciding that any remaining interruption is acceptable, then confirm readiness and inspect the affected job IDs.
Update a drained managed worker
set +xcurl --fail --silent --show-error --proto '=https' --tlsv1.2 \  --output install-cpuos.sh https://cpuos.si/install-cpuos.sh# Inspect the downloaded script before the next command.# Run only after the maintenance decision above.sudo bash ./install-cpuos.sh --updatesudo systemctl status cpuos-worker --no-pagersudo journalctl -u cpuos-worker -n 50 --no-pager

--update preserves the configured control-plane URL and token, verifies the downloaded worker against the release checksum, and restarts the service. Do not pass a new --url or --token-file with update. A stop or restart can interrupt an active job; successful installation is not evidence that the prior job completed.

After interruption, inspect job status and any effects before retrying. If an original submission response was lost, reuse its original Idempotency-Key and identical payload to recover the same job, rather than create another execution. A deliberate new attempt needs a new intent/key. The Python API guide and JavaScript API guide show submission and polling clients.

Reconnect after token revocation without printing secrets

A rejected or revoked node token causes worker authentication exit 77. The managed unit prevents automatic restart on that exit. Create a new node in the same workspace and use its new node token with --reconnect. A workspace API key authorizes application jobs; it cannot replace a node token.

Replace a revoked node credential
set +xcurl --fail --silent --show-error --proto '=https' --tlsv1.2 \  --output install-cpuos.sh https://cpuos.si/install-cpuos.sh# Inspect install-cpuos.sh before continuing. Use Bash for the masked prompt.read -r -s -p "New cpuOS node token: " CPUOS_NODE_TOKENprintf '\n'export CPUOS_NODE_TOKENsudo --preserve-env=CPUOS_NODE_TOKEN bash ./install-cpuos.sh --reconnectunset CPUOS_NODE_TOKENsudo systemctl status cpuos-worker --no-pager

The installer accepts the secret from the masked environment prompt, not as a command-line token argument. Reconnect replaces the credential and preserves the existing control-plane URL unless you explicitly change it. Keep shell tracing off; never paste the credential into code, logs or a support ticket.

Retire the old process on any other host before reusing an identity. Prefer a separate node enrollment for a separate host. A restarted worker cleans interrupted containers carrying its own node identity; it is not a general Docker cleanup tool. Do not run broad container-removal commands to recover unrelated work.

Verify recovery with an intentional small job

When Nodes shows a fresh, Docker-ready worker, submit one new trusted test attempt from the dashboard, such as Python print(2 + 3). Confirm terminal completed status, exit code 0, stdout 5 and no truncation before reopening application traffic. Keep the job ID with your maintenance record.

Then use a representative bounded workload at your normal limits. Check its result, memory-related failures and duration against your application's requirements. A passing small result verifies the connection and execution path; it establishes no throughput or uptime guarantee.

For a broader isolation requirement, read Firecracker vs Docker for AI agents. Current cpuOS jobs use restricted Docker containers on a shared host kernel for trusted team code. Browser execution, arbitrary dependency installation, published SDKs and hostile multi-tenant isolation are outside this pilot.

Questions

Why is the systemd service active while my node is offline?
The process can be alive while runtime images prepare, Docker is unavailable or its control-plane heartbeat is stale. Online requires an unrevoked token, Docker availability and a heartbeat newer than 40 seconds. Inspect the journal and the dashboard before restarting.
Can I run two worker processes with one node token?
Use one process per token across all hosts. Local duplicate processes are rejected, but that does not coordinate two machines. Enroll separate nodes for separate workers; the current worker executes one job at a time.
Does updating the worker wait for its current job?
No. --update restarts the service and can interrupt an active job. Pause submissions in your applications, inspect outstanding attempts and choose whether to wait or cancel before maintenance. Lost jobs are not automatically rerun.
Why did my service stop after I revoked its node token?
Authentication failure exits with status 77, which the managed unit excludes from automatic restart. Create a new node token and use --reconnect through the masked prompt. Repeating a job submission or providing a workspace API key does not repair node authentication.

Related

Connect a worker and run a job

Start with a small trusted Python or Node task, fixed limits and an expected result.

gpuOS · where models think

Need the model too? Run it on gpuOS

gpuOS serves open models on your own GPUs behind one OpenAI-compatible API. Your application can submit authorized actions to cpuOS jobs.