Know what you are operating
A cpuOS worker executes bounded Python and Node jobs on a machine you allocate. The hosted control plane accepts work, stores code and output, and tracks job state. This guide starts after the quickstart: it covers readiness checks, queue diagnosis, maintenance and credential replacement.
Scroll horizontally to see every column.
| Managed worker requirement | Operational consequence |
|---|---|
| Linux with a running systemd manager | The managed service is cpuos-worker; this procedure is for that Linux installation. |
| Node.js 24 or later, installed system-wide | Use /usr/bin/node, /usr/local/bin/node or /bin/node. Home-directory runtime managers are not accepted by the installer. |
| Rootful Docker with Linux containers | The installer requires a docker-group-owned /var/run/docker.sock and checks access as the service account. |
| flock from util-linux and the installer prerequisites | The worker rejects local duplicate processes for the same node token; the installer checks required commands before changing files. |
| One process per node token across all hosts | Enroll a separate node for another worker. A local duplicate check is not a cross-host coordinator. |
| One running job per worker | More host CPUs do not make this worker execute multiple jobs in parallel. |
The service account belongs to the Docker group. Docker daemon access grants root-level control of the host, even though submitted containers run with restrictions. Use a purpose-allocated machine and trusted team code. Do not open a remote Docker API or make the socket world-writable to repair access. Docker Engine security.
The control plane is EU-hosted; the worker runs wherever you place it. Code and results travel through that hosted service. Connecting your own worker does not self-host the control plane, create an air-gapped system or provide Firecracker isolation.
Start with read-only service and Docker checks
Run these commands on the managed Linux worker host. They report status without stopping the service, running a job or printing its credential file. A nonzero status from a failed service is evidence to inspect, not a reason to restart immediately.
sudo systemctl status cpuos-worker --no-pagersudo systemctl show cpuos-worker \ --property=ActiveState --property=SubState --property=Result \ --property=ExecMainStatus --property=NRestartssudo journalctl -u cpuos-worker -n 50 --no-pagercommand -v nodenode --versioncommand -v flocksudo stat --format '%A %U %G %n' /var/run/docker.socksudo runuser -u cpuos-worker -- docker \ --host unix:///var/run/docker.sock info \ --format 'OS={{.OSType}} CPUs={{.NCPU}} MemoryBytes={{.MemTotal}}'Expect Docker's OS to be linux and its daemon to be reachable by cpuos-worker. Compare the Node binary path with the supported system-wide locations. Your interactive shell can resolve a different Node version from the service, so a working home-directory installation is not sufficient. Node.js 24 documentation.
Check Nodes in the dashboard as well as systemd. An active service process can still be preparing images or failing to communicate with the control plane. Initial runtime-image pulls need outbound host connectivity; job containers themselves have networking disabled.
Do not print /etc/cpuos/worker.env, process environments or shell traces in a support report. The root-owned 0600 file contains the node token. Review logs before sharing: hostnames and job IDs can identify your infrastructure and work.
Separate process health, readiness and job eligibility
Scroll horizontally to see every column.
| Observation | What it means | What it does not prove |
|---|---|---|
| Service active, runtime images preparing | The worker can report preparation heartbeats while pulling images; it does not accept a job during that phase. | That Docker is ready to execute work. |
| Node online | Its token is not revoked, Docker is reported available and its heartbeat is newer than 40 seconds. | A reserved execution slot, available free RAM, guaranteed latency or an availability SLA. |
| Node offline | The online check failed, for example due to a stale heartbeat or unavailable Docker. | That code never started, or that a restart is safe for an active attempt. |
| Node revoked | The credential no longer authorizes worker polling. | That resubmitting the same job will repair the worker connection. |
Job admission also compares requested CPU and memory with the capacity reported by an online worker. That report is based on host and Docker capacity, not a continuously measured free-memory budget. A machine already under memory pressure can pass the eligibility check and still fail to execute the workload.
Use node readiness to decide whether submission is possible, then inspect job lifecycle and host measurements to decide whether this workload can run reliably. The pilot supplies neither a liveness promise nor a measured jobs-per-second commitment.
Plan resource limits and queue behavior separately
Scroll horizontally to see every column.
| Current contract | Planning rule |
|---|---|
| 1 or 2 CPUs; 128–2048 MiB per job | Start with 1 CPU and 256 MiB. Increase only after checking realistic inputs and host pressure. |
| 1–120 second execution deadline; default 30 seconds | This bounds an execution attempt, not end-to-end time including queueing and image preparation. |
| One active job per worker | Budget queued work using measured job durations and preparation time, rather than multiplying CPU count by a concurrency assumption. |
| 50 queued jobs per workspace | A full queue rejects new submissions with queue_full; it is not an autoscaling signal that cpuOS will act on. |
| Queued attempts expire after 10 minutes | Inspect queue_expired results and decide whether another intended execution is appropriate. |
| 60 API requests or submissions per minute per workspace | Coordinate submission and polling clients. Creating another API key does not create a separate capacity bucket. |
Docker applies CPU and memory limits to job containers. The worker also disables swap for the configured container memory budget. Resource limits can constrain consumption; they do not reserve unused host RAM or make a memory-heavy transformation fit. Inspect Docker/kernel diagnostics and test the input sizes you intend to process. Docker resource constraints.
cpuOS usage totals multiply worker-reported duration by the requested CPU and memory allocations. They are allocation-based accounting values, not measurements of actual host CPU utilization or resident memory. Keep host monitoring separate when sizing a worker.
For workloads to measure, use the CSV analysis recipe and JSON validation recipe. Vary input size while checking result correctness, duration and failure behavior. A small synthetic example establishes a baseline, not production capacity.
Use the symptom to choose the next check
Scroll horizontally to see every column.
| Symptom | Check | Next action |
|---|---|---|
| Service active but node not ready | Recent worker journal, preparation messages, image-pull connectivity and Docker daemon access. | Resolve the observed prerequisite or connection failure; wait for image preparation rather than sending repeated jobs. |
| 503 no_online_nodes on submission | Node heartbeat, Docker availability and reported CPU/memory against the requested limits. | Restore an eligible worker or choose a smaller request only if its measured workload fits the smaller limits. |
| Job remains queued | Other active jobs, eligible worker capacity, heartbeat and the job's creation time. | Wait within your own budget, reduce incoming submissions, or cancel the particular queued job. Do not assume a second copy will start sooner. |
| 429 queue_full or rate_limit_exceeded | The structured error code, queued count, and how often applications submit or poll. | Slow the relevant clients. Cancel unnecessary queued work intentionally; extra keys do not bypass workspace limits. |
| Failed result with worker_lost | Job ID and result, worker service journal, host or network interruption around the attempt. | Establish whether effects need review before creating a new attempt. Lost jobs are not automatically replayed. |
| Deadline exceeded or nonzero exit code | Terminal error, stderr, outputTruncated, input size, allocated memory and host memory pressure. | Correct the observed failure and re-evaluate limits. An exit code alone is not proof of an out-of-memory cause. |
| Worker exited with status 77 | Authentication/revocation message in its journal and node state in the workspace. | Enroll a new node and reconnect with its token. The managed service deliberately does not restart-loop on this exit. |
| Duplicate-worker or configuration failure | A manual worker and the managed service sharing a node token, or an unsupported prerequisite. | Choose one process for that node identity after reviewing active work. Do not remove locks while another worker may still run. |
For a support report, collect timestamps, the job ID, structured error code, worker version and the relevant non-secret journal excerpt. Include whether the attempt was cold or warm and whether a maintenance action occurred. Redact submitted code and outputs when they contain private data.
Drain application traffic before an update
- Stop new submissions in the applications and team workflow that feed this workspace. There is no drain API or maintenance switch in the current pilot.
- Inspect queued and running jobs in the dashboard, or use authenticated GET /v1/jobs and GET /v1/jobs/:id. Follow list pagination when present so a first page does not hide an older active attempt.
- Choose a maintenance wait budget. Wait for completed, failed or cancelled status and inspect the attempt's result. A running cancellation is asynchronous: verify worker acknowledgement/result or loss handling before treating the worker as idle.
- If you decide to cancel a particular queued or running attempt, use its actual ID with DELETE /v1/jobs/:id. Cancellation does not automatically submit a replacement.
- Download and inspect the current installer. Run --update only after deciding that any remaining interruption is acceptable, then confirm readiness and inspect the affected job IDs.
set +xcurl --fail --silent --show-error --proto '=https' --tlsv1.2 \ --output install-cpuos.sh https://cpuos.si/install-cpuos.sh# Inspect the downloaded script before the next command.# Run only after the maintenance decision above.sudo bash ./install-cpuos.sh --updatesudo systemctl status cpuos-worker --no-pagersudo journalctl -u cpuos-worker -n 50 --no-pager--update preserves the configured control-plane URL and token, verifies the downloaded worker against the release checksum, and restarts the service. Do not pass a new --url or --token-file with update. A stop or restart can interrupt an active job; successful installation is not evidence that the prior job completed.
After interruption, inspect job status and any effects before retrying. If an original submission response was lost, reuse its original Idempotency-Key and identical payload to recover the same job, rather than create another execution. A deliberate new attempt needs a new intent/key. The Python API guide and JavaScript API guide show submission and polling clients.
Reconnect after token revocation without printing secrets
A rejected or revoked node token causes worker authentication exit 77. The managed unit prevents automatic restart on that exit. Create a new node in the same workspace and use its new node token with --reconnect. A workspace API key authorizes application jobs; it cannot replace a node token.
set +xcurl --fail --silent --show-error --proto '=https' --tlsv1.2 \ --output install-cpuos.sh https://cpuos.si/install-cpuos.sh# Inspect install-cpuos.sh before continuing. Use Bash for the masked prompt.read -r -s -p "New cpuOS node token: " CPUOS_NODE_TOKENprintf '\n'export CPUOS_NODE_TOKENsudo --preserve-env=CPUOS_NODE_TOKEN bash ./install-cpuos.sh --reconnectunset CPUOS_NODE_TOKENsudo systemctl status cpuos-worker --no-pagerThe installer accepts the secret from the masked environment prompt, not as a command-line token argument. Reconnect replaces the credential and preserves the existing control-plane URL unless you explicitly change it. Keep shell tracing off; never paste the credential into code, logs or a support ticket.
Retire the old process on any other host before reusing an identity. Prefer a separate node enrollment for a separate host. A restarted worker cleans interrupted containers carrying its own node identity; it is not a general Docker cleanup tool. Do not run broad container-removal commands to recover unrelated work.
Verify recovery with an intentional small job
When Nodes shows a fresh, Docker-ready worker, submit one new trusted test attempt from the dashboard, such as Python print(2 + 3). Confirm terminal completed status, exit code 0, stdout 5 and no truncation before reopening application traffic. Keep the job ID with your maintenance record.
Then use a representative bounded workload at your normal limits. Check its result, memory-related failures and duration against your application's requirements. A passing small result verifies the connection and execution path; it establishes no throughput or uptime guarantee.
For a broader isolation requirement, read Firecracker vs Docker for AI agents. Current cpuOS jobs use restricted Docker containers on a shared host kernel for trusted team code. Browser execution, arbitrary dependency installation, published SDKs and hostile multi-tenant isolation are outside this pilot.