gpuq: document the scheduler (design, contract, ops, failure modes, alternatives)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Nils
2026-07-14 11:54:38 +02:00
co-authored by Claude Fable 5
parent 0252641701
commit 471cf336d0
+146
View File
@@ -0,0 +1,146 @@
# gpuq — bucket-backed multi-node GPU job queue
A minimal scheduler for keeping N GPUs across M hosts busy when the hosts
cannot reach each other directly (rented vast.ai nodes) but all share one
S3 bucket (rclone remote `jspace`, Hetzner). Three shell scripts, no
daemon, no database, no cross-node networking.
```
scripts/gpuq_worker.sh one per GPU runs jobs from its own queue
scripts/gpuq_supervisor.sh one per node restarts workers, heartbeats
scripts/gpuq_submit.sh anywhere enqueue jobs / show status
```
## Design in one paragraph
Every GPU gets a private queue directory in the bucket:
`jspace:jspace/gpuq/<worker-id>/pending/`. A worker loop polls its own
directory (20 s), pulls job files, runs them oldest-first-by-name with
`CUDA_VISIBLE_DEVICES` pinned, and moves job + full log to `done/`.
There is **no work stealing and no job claiming** — each queue has exactly
one consumer, so the S3 lack of atomic move can never double-run a job.
Load balancing happens at submit time (you pick the worker); ordering
happens in the filename (`00_`, `01_`, …). The bucket is both transport
and durable state: a dead node's pending jobs stay visible and can be
moved to another worker with one `rclone move`.
## Bucket layout
```
gpuq/
_health/<node-id>.txt heartbeat, rewritten every 60 s
<node-id>-gpu<i>/
pending/<job>.sh waiting (consumed oldest-first by name)
done/<job>.sh finished job file
done/<job>.log its full stdout+stderr, ends with "rc=N"
```
Local mirror on the node: `~/gpuq/<worker-id>/` (pending copy, logs).
## Node setup (once per node, after git clone/pull of ~/jspace)
Requirements: `rclone` configured with the `jspace` remote (see
`.s3-credentials.txt` on the Spark — do NOT commit it), tmux session
`ssh_tmux`, GPUs visible to `nvidia-smi`.
Start ONLY the supervisor — it spawns and maintains all workers itself:
```bash
tmux new-window -t ssh_tmux -n gpuq_sup \
'bash ~/jspace/scripts/gpuq_supervisor.sh node1 5' # node-id, #GPUs
```
The supervisor, every 60 s:
1. respawns any missing `gpuq_worker.sh node1-gpu<i>` as tmux window
`gpuq<i>` (nohup fallback if tmux is unreachable);
2. publishes `gpuq/_health/node1.txt`: timestamp, uptime, per-GPU
util/mem, disk on `/` and `/dev/shm`, live worker count, local
pending count.
There is deliberately no supervisor-of-supervisors: a stale heartbeat
(> ~2 min) IS the signal that the node or supervisor died — check it
from outside.
## Job contract
A job is a plain bash file. The worker provides `CUDA_VISIBLE_DEVICES`;
everything else is the job's business. Template:
```bash
#!/bin/bash
cd ~/jspace/scripts
P=~/jspace/.venv/bin/python # /venv/main/bin/python on vast images
export HF_HUB_OFFLINE=1 # model already cached
LOOP_OUT=~/jspace/results-loop-12b $P train_merge_code.py --seed 7 --alpha 0.15
rclone copy ~/jspace/results-loop-12b jspace:jspace/results-12b/ \
--include "adapter_code_s7*" --include "*.json" -q
```
Rules of thumb:
- **Jobs sync their own results** to the bucket at the end (the queue only
ships the job's log automatically, not its artifacts).
- One GPU per job; never set CUDA_VISIBLE_DEVICES yourself.
- Idempotence is on the submitter: if you resubmit a job, it reruns.
- Exit code lands as the last line of the log (`rc=0`).
## Submitting and monitoring (from any machine with the remote)
```bash
# ordered pair on one GPU (train, then eval — order = filename sort):
scripts/gpuq_submit.sh node1-gpu0 00_train_a15.sh 01_eval_a15.sh
# fan five seeds across five GPUs:
for s in 0 1 2 3 4; do
scripts/gpuq_submit.sh node1-gpu$s seed$s.sh
done
scripts/gpuq_submit.sh --status # pending + last 20 finished
rclone cat jspace:jspace/gpuq/_health/node1.txt # node health
rclone cat jspace:jspace/gpuq/node1-gpu0/done/00_train_a15.log | tail -50
```
## Dependencies
- Same-GPU sequential dependency: filename prefixes on one worker
(`00_train.sh`, `01_eval.sh`). This covers nearly all our chains.
- Cross-GPU/cross-node dependencies: not supported by design. Either
submit the dependent job after seeing the first finish in `--status`,
or make the job itself poll the bucket for its input artifact before
starting (a `until rclone lsf ...; do sleep 60; done` preamble).
## Failure modes and recovery
| failure | effect | recovery |
|---|---|---|
| job crashes (rc≠0) | worker moves on to next job | read `done/<job>.log`; resubmit fixed job |
| worker dies | its queue stalls | supervisor respawns it ≤ 60 s |
| supervisor dies | workers keep running; heartbeat goes stale | restart supervisor window |
| node dies | heartbeat stale; pending jobs preserved in bucket | `rclone move gpuq/node1-gpu0/pending gpuq/node2-gpu3/pending` |
| duplicate submit | job runs twice (queues are dumb) | submitter's responsibility |
| partial upload read | impossible — S3 uploads are atomic (objects appear only complete) | — |
Local `~/gpuq/` is scratch; the bucket is the source of truth.
## Why not X (considered alternatives)
- **Slurm/K8s**: assume stable, mutually-reachable nodes; wrong shape for
ephemeral single-tenant rentals.
- **Ray / SkyPilot / dstack**: viable (dstack has native vast.ai support
and is the thing to try for the *next* provisioning burst), but they
own provisioning; gpuq feeds nodes that already exist.
- **iroh p2p binary**: would give push dispatch, live log streaming, and
direct node↔node transfer through NAT (QUIC hole-punching) — attractive
if the fleet becomes permanent. Rejected for now: jobs run 10 min2 h,
so poll latency is irrelevant, and durable queue state in S3 comes free
vs. a coordinator protocol we'd have to write and debug. For ad-hoc live
log streaming, n0's prebuilt `dumbpipe` works without writing code.
- **Marker-file watchers / pkill chains** (our previous approach): see
LESSONS.md — self-matching kill patterns and watcher pileups burned us
repeatedly. gpuq replaces markers with per-consumer queues and replaces
pkill with the supervisor owning worker lifecycle.
## Security notes
The bucket credentials on rented nodes are readable by the host operator:
keep the `jspace` remote scoped to this project's bucket, and rotate the
key after each rental burst (already flagged in vast-ai-notes.md).