cpuos

Tutorials · 7 min read · updated Oct 7, 2026

Batch Python jobs with bounded inline inputs and checked results

Split small inline datasets into bounded Python jobs with cpuOS. Plan chunks in your application, control queue pressure and merge only validated results.

The current pilot runs trusted Python and Node jobs on your Docker worker. Containers share its kernel. Browser and repository workflows need capabilities beyond this pilot.

On this page

Partition only work that is independently computable

Batch Python execution is useful when one application task contains several small independent calculations. Your application partitions the input, submits one bounded cpuOS job per chunk and merges checked results. cpuOS provides individual jobs, not a native batch upload, distributed data engine or automatic fan-out/fan-in workflow. Keep the manifest and merge policy in your application.

This recipe sums squared synthetic integers. Each row contributes independently, so adding chunk totals reproduces the full total. Global sorting, deduplication, averages and joins need different merge rules: averaging chunk averages, for example, is wrong when chunk sizes differ. Define the global invariant before choosing a partition size.

Keep a manifest of stable row and batch identities

Scroll horizontally to see every column.

Application recordPurpose
Run identifierOne business task across all chunks and approved attempts.
Batch ID and expected row IDsProve the returned result belongs to the planned chunk.
Exact source and requested limitsRecover the same intent and audit the submitted transformation.
Idempotency-Key and job IDRecover ambiguous submission without creating an extra execution.
Terminal result and merge stateTrack failed chunks and prevent adding a successful chunk twice.

Validate unique row IDs across the complete input before partitioning. Checking uniqueness inside each job cannot catch a duplicate split across different chunks. The worker program repeats per-chunk checks because encoded data still needs its own contract. Generate a distinct run identifier for a new business task, and persist the complete manifest before scheduling its first job.

Run one complete standard-library batch fixture

batch.py: three inline synthetic rows, Python 3.13
import jsonINPUT_JSON = "{\"batch_id\":\"fixture-01:001\",\"rows\":[{\"id\":\"row-001\",\"value\":3},{\"id\":\"row-002\",\"value\":4},{\"id\":\"row-003\",\"value\":5}]}"import reimport systry:    payload = json.loads(INPUT_JSON)    if type(payload) is not dict or set(payload) != {"batch_id", "rows"}:        raise ValueError("Invalid batch fields.")    batch_id = payload["batch_id"]    if not isinstance(batch_id, str) or not re.fullmatch(r"[A-Za-z0-9:_-]{1,64}", batch_id):        raise ValueError("Invalid batch ID.")    rows = payload["rows"]    if type(rows) is not list or not 1 <= len(rows) <= 25:        raise ValueError("Expected 1 to 25 rows.")    seen = set()    total = 0    for row in rows:        if type(row) is not dict or set(row) != {"id", "value"}:            raise ValueError("Invalid row fields.")        row_id = row["id"]        if not isinstance(row_id, str) or not re.fullmatch(r"row-[0-9]{3}", row_id) or row_id in seen:            raise ValueError("Invalid or duplicate row ID.")        value = row["value"]        if type(value) is not int or not 0 <= value <= 1000:            raise ValueError("Expected a bounded integer value.")        seen.add(row_id)        total += value * value    print(json.dumps({        "batch_id": batch_id, "count": len(rows),        "row_ids": sorted(seen), "sum_squares": total,    }, allow_nan=False, sort_keys=True))except ValueError as error:    print(str(error), file=sys.stderr)    sys.exit(1)
Expected complete parsed result
{  "batch_id": "fixture-01:001",  "count": 3,  "row_ids": [    "row-001",    "row-002",    "row-003"  ],  "sum_squares": 50}

Run python3 batch.py locally and compare the complete object. Values three, four and five contribute 9, 16 and 25, giving total 50 across three distinct rows. The program accepts one to twenty-five records, bounded integer values and an explicit batch identifier. It emits no partial result when input validation fails.

The input is a JSON string decoded with Python json.loads, not a sequence of interpolated Python expressions. Prices, labels or other future fields need their own validation rules. Keep submitted datasets small and minimize sensitive values: the hosted control plane receives the source containing this inline input.

Prepare chunks and measure the complete source in the application

plan-batches.mjs: build two job request bodies locally with Node 24
// Run in your application with Node 24. This prepares jobs; it does not submit them.const program = "import re\nimport sys\n\ntry:\n    payload = json.loads(INPUT_JSON)\n    if type(payload) is not dict or set(payload) != {\"batch_id\", \"rows\"}:\n        raise ValueError(\"Invalid batch fields.\")\n    batch_id = payload[\"batch_id\"]\n    if not isinstance(batch_id, str) or not re.fullmatch(r\"[A-Za-z0-9:_-]{1,64}\", batch_id):\n        raise ValueError(\"Invalid batch ID.\")\n    rows = payload[\"rows\"]\n    if type(rows) is not list or not 1 <= len(rows) <= 25:\n        raise ValueError(\"Expected 1 to 25 rows.\")\n    seen = set()\n    total = 0\n    for row in rows:\n        if type(row) is not dict or set(row) != {\"id\", \"value\"}:\n            raise ValueError(\"Invalid row fields.\")\n        row_id = row[\"id\"]\n        if not isinstance(row_id, str) or not re.fullmatch(r\"row-[0-9]{3}\", row_id) or row_id in seen:\n            raise ValueError(\"Invalid or duplicate row ID.\")\n        value = row[\"value\"]\n        if type(value) is not int or not 0 <= value <= 1000:\n            raise ValueError(\"Expected a bounded integer value.\")\n        seen.add(row_id)\n        total += value * value\n    print(json.dumps({\n        \"batch_id\": batch_id, \"count\": len(rows),\n        \"row_ids\": sorted(seen), \"sum_squares\": total,\n    }, allow_nan=False, sort_keys=True))\nexcept ValueError as error:\n    print(str(error), file=sys.stderr)\n    sys.exit(1)"const rows = [{"id":"row-001","value":3},{"id":"row-002","value":4},{"id":"row-003","value":5}]const runId = "fixture-01"const size = 2  // Demonstration; the Python recipe accepts at most 25 rows per job.const requests = []const ids = new Set()for (const row of rows) {  if (!row || Object.keys(row).length !== 2 || !Object.hasOwn(row, "id")      || !Object.hasOwn(row, "value") || typeof row.id !== "string"      || row.id.length !== 7 || !/^row-[0-9]{3}$/.test(row.id) || ids.has(row.id)      || !Number.isInteger(row.value) || row.value < 0 || row.value > 1000) {    throw new Error("Invalid or duplicate source row.")  }  ids.add(row.id)}for (let offset = 0; offset < rows.length; offset += size) {  const number = String(requests.length + 1).padStart(3, "0")  const batchId = runId + ":" + number  const payload = { batch_id: batchId, rows: rows.slice(offset, offset + size) }  // Encode JSON as a Python string literal, never interpolate raw data as code.  const code = "import json\nINPUT_JSON = " + JSON.stringify(JSON.stringify(payload))    + "\n" + program  if (Buffer.byteLength(code, "utf8") > 64 * 1024) {    throw new Error("Reduce the chunk size; complete source exceeds 64 KiB.")  }  requests.push({ name: batchId, template: "python", code,    vcpu: 1, memoryMb: 256, timeoutSeconds: 30 })}console.log(JSON.stringify(requests.map((request) => ({  name: request.name, template: request.template,}))))// Expected: two request descriptors, fixture-01:001 and fixture-01:002.// Persist each request with its own Idempotency-Key before calling POST /v1/jobs.

Run node plan-batches.mjs on the submitting machine. It prepares two requests for the three rows, splitting them into groups of two and one. It does not call the jobs API. Send each prepared body through the existing HTTP job client contract with its own persisted intent key, or adapt the Python client to your application scheduler.

The complete code string, including the Python program and encoded JSON, must fit the 64-KiB UTF-8 source bound. Count bytes with Node Buffer.byteLength, rather than JavaScript character length. A row-count bound is useful but cannot replace a byte bound when record sizes vary.

The demonstration's fixed chunk size is an application choice, not a cpuOS API field or a measured performance recommendation. For a variable-sized input, decrease the chunk size when source would exceed its budget. Reject any single record that cannot fit; do not truncate input silently. No input-file upload, package installation or job network is involved.

Use a bounded submission window instead of flooding the queue

Start with sequential execution, then increase the number of in-flight jobs only after checking worker behavior and the application's request budget. The workspace currently allows fifty queued jobs and shares API rate limits across keys. A smaller application window makes capacity failures, cancellation and recovery easier to inspect. It does not reserve CPU resources or promise simultaneous starts.

Apply one consistent resource policy to each chunk, such as the illustrated one CPU, 256 MiB and thirty-second timeout, then test that policy with representative inputs. Those values are starting parameters, not throughput measurements. If a chunk cannot fit the pilot's maximum two CPUs, 2048 MiB or 120-second deadline, change the algorithm or split it further.

On queue_full, pause new submissions and inspect accepted jobs. On no_online_nodes, restore an eligible worker before continuing. Use CPU resource planning and worker operations to distinguish request bounds from host readiness.

Merge only a complete set of validated batch results

  • Require completed status, exitCode zero, no execution error and outputTruncated false for every chunk.
  • Parse the entire stdout JSON and check exactly the batch_id, count, row_ids and sum_squares fields.
  • Compare the batch ID and sorted row IDs with that chunk's manifest; require count to equal the planned row count.
  • Check integer types and total bounds, then add each accepted chunk exactly once.
  • Require the union of accepted row IDs to match the complete input manifest before marking the run complete.

The two planned fixture chunks contribute totals 25 and 25, with counts two and one. Their union covers the three source IDs, and their total 50 matches the single-job fixture. Store each validated result atomically with its merge state. Completion order should not affect the answer, and receiving the same job result twice should not add the total twice.

A failed chunk leaves the application run incomplete. Report that outcome explicitly or obtain an approved new attempt for that chunk. Recover an ambiguous POST with the original key; request a new computation with a new key after reviewing the cause. Structured result validation and asynchronous lifecycle handling provide the two checks this merge depends on.

Measure the model step and the CPU batch step separately

An agent may propose the transformation or explain the checked aggregate, while your application performs partitioning and job control. The model should receive a compact final result and explicit missing-chunk failures, rather than raw output from every attempt. gpuOS model benchmarking covers inference measurements; CPU queue delay and script execution need their own observations. There is no automatic built-in connection between the two products.

Questions

Does cpuOS offer a batch file-upload API?
No. This recipe creates individual Python jobs with small inline JSON inputs. Your application partitions data, persists a manifest, submits jobs and merges results; file transfer and package installation are unavailable.
How do I avoid double-counting retried batch jobs?
Persist one intent key per unchanged submission and one merge record per logical batch. Recover an uncertain request with the same key, then merge a validated successful batch exactly once. A new approved computation attempt needs a new key.
Can I merge whichever chunks finished and call the batch successful?
Only if partial completion is an explicit different application contract. This recipe requires every planned row ID exactly once and every chunk's checked terminal result. Any missing or failed chunk leaves the run incomplete.

Related

Connect a worker and run a job

Start with a small trusted Python or Node task, fixed limits and an expected result.

gpuOS · where models think

Need the model too? Run it on gpuOS

gpuOS serves open models on your own GPUs behind one OpenAI-compatible API. Your application can submit authorized actions to cpuOS jobs.