147 lines
6.2 KiB
Markdown
147 lines
6.2 KiB
Markdown
# gpuq — bucket-backed multi-node GPU job queue
|
||
|
||
A minimal scheduler for keeping N GPUs across M hosts busy when the hosts
|
||
cannot reach each other directly (rented vast.ai nodes) but all share one
|
||
S3 bucket (rclone remote `jspace`, Hetzner). Three shell scripts, no
|
||
daemon, no database, no cross-node networking.
|
||
|
||
```
|
||
scripts/gpuq_worker.sh one per GPU runs jobs from its own queue
|
||
scripts/gpuq_supervisor.sh one per node restarts workers, heartbeats
|
||
scripts/gpuq_submit.sh anywhere enqueue jobs / show status
|
||
```
|
||
|
||
## Design in one paragraph
|
||
|
||
Every GPU gets a private queue directory in the bucket:
|
||
`jspace:jspace/gpuq/<worker-id>/pending/`. A worker loop polls its own
|
||
directory (20 s), pulls job files, runs them oldest-first-by-name with
|
||
`CUDA_VISIBLE_DEVICES` pinned, and moves job + full log to `done/`.
|
||
There is **no work stealing and no job claiming** — each queue has exactly
|
||
one consumer, so the S3 lack of atomic move can never double-run a job.
|
||
Load balancing happens at submit time (you pick the worker); ordering
|
||
happens in the filename (`00_`, `01_`, …). The bucket is both transport
|
||
and durable state: a dead node's pending jobs stay visible and can be
|
||
moved to another worker with one `rclone move`.
|
||
|
||
## Bucket layout
|
||
|
||
```
|
||
gpuq/
|
||
_health/<node-id>.txt heartbeat, rewritten every 60 s
|
||
<node-id>-gpu<i>/
|
||
pending/<job>.sh waiting (consumed oldest-first by name)
|
||
done/<job>.sh finished job file
|
||
done/<job>.log its full stdout+stderr, ends with "rc=N"
|
||
```
|
||
|
||
Local mirror on the node: `~/gpuq/<worker-id>/` (pending copy, logs).
|
||
|
||
## Node setup (once per node, after git clone/pull of ~/jspace)
|
||
|
||
Requirements: `rclone` configured with the `jspace` remote (see
|
||
`.s3-credentials.txt` on the Spark — do NOT commit it), tmux session
|
||
`ssh_tmux`, GPUs visible to `nvidia-smi`.
|
||
|
||
Start ONLY the supervisor — it spawns and maintains all workers itself:
|
||
|
||
```bash
|
||
tmux new-window -t ssh_tmux -n gpuq_sup \
|
||
'bash ~/jspace/scripts/gpuq_supervisor.sh node1 5' # node-id, #GPUs
|
||
```
|
||
|
||
The supervisor, every 60 s:
|
||
1. respawns any missing `gpuq_worker.sh node1-gpu<i>` as tmux window
|
||
`gpuq<i>` (nohup fallback if tmux is unreachable);
|
||
2. publishes `gpuq/_health/node1.txt`: timestamp, uptime, per-GPU
|
||
util/mem, disk on `/` and `/dev/shm`, live worker count, local
|
||
pending count.
|
||
|
||
There is deliberately no supervisor-of-supervisors: a stale heartbeat
|
||
(> ~2 min) IS the signal that the node or supervisor died — check it
|
||
from outside.
|
||
|
||
## Job contract
|
||
|
||
A job is a plain bash file. The worker provides `CUDA_VISIBLE_DEVICES`;
|
||
everything else is the job's business. Template:
|
||
|
||
```bash
|
||
#!/bin/bash
|
||
cd ~/jspace/scripts
|
||
P=~/jspace/.venv/bin/python # /venv/main/bin/python on vast images
|
||
export HF_HUB_OFFLINE=1 # model already cached
|
||
LOOP_OUT=~/jspace/results-loop-12b $P train_merge_code.py --seed 7 --alpha 0.15
|
||
rclone copy ~/jspace/results-loop-12b jspace:jspace/results-12b/ \
|
||
--include "adapter_code_s7*" --include "*.json" -q
|
||
```
|
||
|
||
Rules of thumb:
|
||
- **Jobs sync their own results** to the bucket at the end (the queue only
|
||
ships the job's log automatically, not its artifacts).
|
||
- One GPU per job; never set CUDA_VISIBLE_DEVICES yourself.
|
||
- Idempotence is on the submitter: if you resubmit a job, it reruns.
|
||
- Exit code lands as the last line of the log (`rc=0`).
|
||
|
||
## Submitting and monitoring (from any machine with the remote)
|
||
|
||
```bash
|
||
# ordered pair on one GPU (train, then eval — order = filename sort):
|
||
scripts/gpuq_submit.sh node1-gpu0 00_train_a15.sh 01_eval_a15.sh
|
||
|
||
# fan five seeds across five GPUs:
|
||
for s in 0 1 2 3 4; do
|
||
scripts/gpuq_submit.sh node1-gpu$s seed$s.sh
|
||
done
|
||
|
||
scripts/gpuq_submit.sh --status # pending + last 20 finished
|
||
rclone cat jspace:jspace/gpuq/_health/node1.txt # node health
|
||
rclone cat jspace:jspace/gpuq/node1-gpu0/done/00_train_a15.log | tail -50
|
||
```
|
||
|
||
## Dependencies
|
||
|
||
- Same-GPU sequential dependency: filename prefixes on one worker
|
||
(`00_train.sh`, `01_eval.sh`). This covers nearly all our chains.
|
||
- Cross-GPU/cross-node dependencies: not supported by design. Either
|
||
submit the dependent job after seeing the first finish in `--status`,
|
||
or make the job itself poll the bucket for its input artifact before
|
||
starting (a `until rclone lsf ...; do sleep 60; done` preamble).
|
||
|
||
## Failure modes and recovery
|
||
|
||
| failure | effect | recovery |
|
||
|---|---|---|
|
||
| job crashes (rc≠0) | worker moves on to next job | read `done/<job>.log`; resubmit fixed job |
|
||
| worker dies | its queue stalls | supervisor respawns it ≤ 60 s |
|
||
| supervisor dies | workers keep running; heartbeat goes stale | restart supervisor window |
|
||
| node dies | heartbeat stale; pending jobs preserved in bucket | `rclone move gpuq/node1-gpu0/pending gpuq/node2-gpu3/pending` |
|
||
| duplicate submit | job runs twice (queues are dumb) | submitter's responsibility |
|
||
| partial upload read | impossible — S3 uploads are atomic (objects appear only complete) | — |
|
||
|
||
Local `~/gpuq/` is scratch; the bucket is the source of truth.
|
||
|
||
## Why not X (considered alternatives)
|
||
|
||
- **Slurm/K8s**: assume stable, mutually-reachable nodes; wrong shape for
|
||
ephemeral single-tenant rentals.
|
||
- **Ray / SkyPilot / dstack**: viable (dstack has native vast.ai support
|
||
and is the thing to try for the *next* provisioning burst), but they
|
||
own provisioning; gpuq feeds nodes that already exist.
|
||
- **iroh p2p binary**: would give push dispatch, live log streaming, and
|
||
direct node↔node transfer through NAT (QUIC hole-punching) — attractive
|
||
if the fleet becomes permanent. Rejected for now: jobs run 10 min–2 h,
|
||
so poll latency is irrelevant, and durable queue state in S3 comes free
|
||
vs. a coordinator protocol we'd have to write and debug. For ad-hoc live
|
||
log streaming, n0's prebuilt `dumbpipe` works without writing code.
|
||
- **Marker-file watchers / pkill chains** (our previous approach): see
|
||
LESSONS.md — self-matching kill patterns and watcher pileups burned us
|
||
repeatedly. gpuq replaces markers with per-consumer queues and replaces
|
||
pkill with the supervisor owning worker lifecycle.
|
||
|
||
## Security notes
|
||
|
||
The bucket credentials on rented nodes are readable by the host operator:
|
||
keep the `jspace` remote scoped to this project's bucket, and rotate the
|
||
key after each rental burst (already flagged in vast-ai-notes.md).
|