228 lines
10 KiB
Markdown
228 lines
10 KiB
Markdown
# gpuq — bucket-backed multi-node GPU job queue
|
||
|
||
A minimal scheduler for keeping N GPUs across M hosts busy when the hosts
|
||
cannot reach each other directly (rented vast.ai nodes) but all share one
|
||
S3 bucket (rclone remote `jspace`, Hetzner). Three shell scripts, no
|
||
daemon, no database, no cross-node networking.
|
||
|
||
```
|
||
scripts/gpuq_worker.sh one per GPU runs jobs from its own queue
|
||
scripts/gpuq_supervisor.sh one per node restarts workers, heartbeats
|
||
scripts/gpuq_submit.sh anywhere enqueue jobs / show status
|
||
scripts/gpuq_dispatcher.sh ONE, always-on assigns class jobs to idle GPUs
|
||
```
|
||
|
||
## Design in one paragraph
|
||
|
||
Every GPU gets a private queue directory in the bucket:
|
||
`jspace:jspace/gpuq/<worker-id>/pending/`. A worker loop polls its own
|
||
directory (20 s), pulls job files, runs them oldest-first-by-name with
|
||
`CUDA_VISIBLE_DEVICES` pinned, and moves job + full log to `done/`.
|
||
There is **no work stealing and no job claiming** — each queue has exactly
|
||
one consumer, so the S3 lack of atomic move can never double-run a job.
|
||
Load balancing happens at submit time (you pick the worker); ordering
|
||
happens in the filename (`00_`, `01_`, …). The bucket is both transport
|
||
and durable state: a dead node's pending jobs stay visible and can be
|
||
moved to another worker with one `rclone move`.
|
||
|
||
## Data contract (selective sync)
|
||
|
||
Jobs declare what they need and what they produce as comments; the worker
|
||
does the rest — no node needs everything, and results survive node death:
|
||
|
||
```
|
||
# gpuq-in: results-loop/mbpp_data.json results-loop/adapter_code.pt
|
||
# gpuq-out: results-loop/eval_code_gate*.json results/jbar_31b.pt
|
||
```
|
||
|
||
Paths are repo-relative, mirrored at `jspace:jspace/<same path>`.
|
||
Inputs are pulled before the job starts. Outputs are pushed every
|
||
`$GPUQ_SYNC_EVERY` seconds (default 120) while the job runs and once at
|
||
exit — a dying node loses at most one interval. The in-flight log is
|
||
also mirrored to `gpuq/<worker>/running/` each cycle. Globs allowed in
|
||
basenames; repeat lines for more paths; jobs with no declarations sync
|
||
nothing (backward compatible). Manual counterpart from any machine:
|
||
|
||
```
|
||
scripts/gpuq_sync.sh pull results-loop "eval_code_gate*.json"
|
||
scripts/gpuq_sync.sh push results-loop "adapter_code.pt"
|
||
```
|
||
|
||
Nodes need the rclone remote configured for any of this; keep a scoped
|
||
(write-limited) key for rented nodes. Without credentials the worker
|
||
still runs jobs — sync lines just no-op.
|
||
|
||
## Bucket layout
|
||
|
||
```
|
||
gpuq/
|
||
_health/<node-id>.txt heartbeat, rewritten every 60 s
|
||
<node-id>-gpu<i>/
|
||
pending/<job>.sh waiting (consumed oldest-first by name)
|
||
done/<job>.sh finished job file
|
||
done/<job>.log its full stdout+stderr, ends with "rc=N"
|
||
```
|
||
|
||
Local mirror on the node: `~/gpuq/<worker-id>/` (pending copy, logs).
|
||
|
||
## Node setup (once per node, after git clone/pull of ~/jspace)
|
||
|
||
Requirements: `rclone` configured with the `jspace` remote (see
|
||
`.s3-credentials.txt` on the Spark — do NOT commit it), tmux session
|
||
`ssh_tmux`, GPUs visible to `nvidia-smi`.
|
||
|
||
Start ONLY the supervisor — it spawns and maintains all workers itself:
|
||
|
||
```bash
|
||
tmux new-window -t ssh_tmux -n gpuq_sup \
|
||
'bash ~/jspace/scripts/gpuq_supervisor.sh node1 5' # node-id, #GPUs
|
||
```
|
||
|
||
The supervisor, every 60 s:
|
||
1. respawns any missing `gpuq_worker.sh node1-gpu<i>` as tmux window
|
||
`gpuq<i>` (nohup fallback if tmux is unreachable);
|
||
2. publishes `gpuq/_health/node1.txt`: timestamp, uptime, per-GPU
|
||
util/mem, disk on `/` and `/dev/shm`, live worker count, local
|
||
pending count.
|
||
|
||
There is deliberately no supervisor-of-supervisors: a stale heartbeat
|
||
(> ~2 min) IS the signal that the node or supervisor died — check it
|
||
from outside.
|
||
|
||
## Job contract
|
||
|
||
A job is a plain bash file. The worker provides `CUDA_VISIBLE_DEVICES`;
|
||
everything else is the job's business. Template:
|
||
|
||
```bash
|
||
#!/bin/bash
|
||
cd ~/jspace/scripts
|
||
P=~/jspace/.venv/bin/python # /venv/main/bin/python on vast images
|
||
export HF_HUB_OFFLINE=1 # model already cached
|
||
LOOP_OUT=~/jspace/results-loop-12b $P train_merge_code.py --seed 7 --alpha 0.15
|
||
rclone copy ~/jspace/results-loop-12b jspace:jspace/results-12b/ \
|
||
--include "adapter_code_s7*" --include "*.json" -q
|
||
```
|
||
|
||
Rules of thumb:
|
||
- **Jobs sync their own results** to the bucket at the end (the queue only
|
||
ships the job's log automatically, not its artifacts).
|
||
- One GPU per job; never set CUDA_VISIBLE_DEVICES yourself.
|
||
- Idempotence is on the submitter: if you resubmit a job, it reruns.
|
||
- Exit code lands as the last line of the log (`rc=0`).
|
||
|
||
## Submitting and monitoring (from any machine with the remote)
|
||
|
||
```bash
|
||
# ordered pair on one GPU (train, then eval — order = filename sort):
|
||
scripts/gpuq_submit.sh node1-gpu0 00_train_a15.sh 01_eval_a15.sh
|
||
|
||
# fan five seeds across five GPUs:
|
||
for s in 0 1 2 3 4; do
|
||
scripts/gpuq_submit.sh node1-gpu$s seed$s.sh
|
||
done
|
||
|
||
scripts/gpuq_submit.sh --status # pending + last 20 finished
|
||
rclone cat jspace:jspace/gpuq/_health/node1.txt # node health
|
||
rclone cat jspace:jspace/gpuq/node1-gpu0/done/00_train_a15.log | tail -50
|
||
```
|
||
|
||
## Dependencies
|
||
|
||
- Same-GPU sequential dependency: filename prefixes on one worker
|
||
(`00_train.sh`, `01_eval.sh`). This covers nearly all our chains.
|
||
- Cross-GPU/cross-node dependencies: not supported by design. Either
|
||
submit the dependent job after seeing the first finish in `--status`,
|
||
or make the job itself poll the bucket for its input artifact before
|
||
starting (a `until rclone lsf ...; do sleep 60; done` preamble).
|
||
|
||
## Failure modes and recovery
|
||
|
||
| failure | effect | recovery |
|
||
|---|---|---|
|
||
| job crashes (rc≠0) | worker moves on to next job | read `done/<job>.log`; resubmit fixed job |
|
||
| worker dies | its queue stalls | supervisor respawns it ≤ 60 s |
|
||
| supervisor dies | workers keep running; heartbeat goes stale | restart supervisor window |
|
||
| node dies | heartbeat stale; pending jobs preserved in bucket | `rclone move gpuq/node1-gpu0/pending gpuq/node2-gpu3/pending` |
|
||
| duplicate submit | job runs twice (queues are dumb) | submitter's responsibility |
|
||
| partial upload read | impossible — S3 uploads are atomic (objects appear only complete) | — |
|
||
|
||
Local `~/gpuq/` is scratch; the bucket is the source of truth.
|
||
|
||
## Why not X (considered alternatives)
|
||
|
||
- **Slurm/K8s**: assume stable, mutually-reachable nodes; wrong shape for
|
||
ephemeral single-tenant rentals.
|
||
- **Ray / SkyPilot / dstack**: viable (dstack has native vast.ai support
|
||
and is the thing to try for the *next* provisioning burst), but they
|
||
own provisioning; gpuq feeds nodes that already exist.
|
||
- **iroh p2p binary**: would give push dispatch, live log streaming, and
|
||
direct node↔node transfer through NAT (QUIC hole-punching) — attractive
|
||
if the fleet becomes permanent. Rejected for now: jobs run 10 min–2 h,
|
||
so poll latency is irrelevant, and durable queue state in S3 comes free
|
||
vs. a coordinator protocol we'd have to write and debug. For ad-hoc live
|
||
log streaming, n0's prebuilt `dumbpipe` works without writing code.
|
||
- **Marker-file watchers / pkill chains** (our previous approach): see
|
||
LESSONS.md — self-matching kill patterns and watcher pileups burned us
|
||
repeatedly. gpuq replaces markers with per-consumer queues and replaces
|
||
pkill with the supervisor owning worker lifecycle.
|
||
|
||
## Security notes
|
||
|
||
The bucket credentials on rented nodes are readable by the host operator:
|
||
keep the `jspace` remote scoped to this project's bucket, and rotate the
|
||
key after each rental burst (already flagged in vast-ai-notes.md).
|
||
|
||
---
|
||
|
||
## As deployed (2026-07-14) — operator cheat sheet
|
||
|
||
**Fleet:**
|
||
|
||
| node | workers | supervisor lives in |
|
||
|---|---|---|
|
||
| `node3` (8×H100, vast) | `node3-gpu0` … `node3-gpu7` | tmux `ssh_tmux:gpuq_sup` |
|
||
| `node2` (2×H100, vast) | `node2-gpu0`, `node2-gpu1` | tmux `ssh_tmux:gpuq_sup` (drain watcher armed on the Spark) |
|
||
| `spark` (GB10) | `spark-gpu0` | detached, tmux session `gpuq_spark` |
|
||
|
||
**Daily usage (from the Spark or anywhere with the `jspace` remote):**
|
||
|
||
```bash
|
||
bash scripts/gpuq_submit.sh --class h100 myjob.sh # PREFERRED: class queue;
|
||
# dispatcher assigns to the
|
||
# next idle GPU in class
|
||
bash scripts/gpuq_submit.sh node3-gpu4 myjob.sh # pin to a specific GPU
|
||
# (ordered chains only)
|
||
bash scripts/gpuq_submit.sh --status # pending + last 20 done
|
||
rclone cat jspace:jspace/gpuq/_health/node3.txt # heartbeat; stale >2 min = trouble
|
||
rclone cat jspace:jspace/gpuq/node3-gpu4/done/myjob.log | tail -30
|
||
```
|
||
|
||
**Job conventions (learned the hard way — see LESSONS.md):**
|
||
|
||
1. First line of work: `cd <repo> && git pull origin main -q` — jobs
|
||
self-update; never rely on what a node cloned at boot (repo is public,
|
||
no keys needed).
|
||
2. `HF_HUB_OFFLINE=1` ONLY if every model/dataset the job touches is already
|
||
cached on that node — it blocks `datasets` downloads too (bit us twice).
|
||
3. Jobs end with their own `rclone copy` of artifacts to the bucket; the
|
||
queue ships only the log automatically.
|
||
4. Env the job must set itself: venv activation (`/venv/main` on vast,
|
||
`~/jspace/.venv` on the Spark), `JLENS_MODEL`, `JLENS_BAND`, `LOOP_OUT`,
|
||
`HF_HOME`. The worker provides exactly one thing: `CUDA_VISIBLE_DEVICES`.
|
||
5. Ordering on one GPU = filename sort (`00_train.sh`, `01_eval.sh`).
|
||
Cross-GPU deps: submit after the prerequisite shows in `--status`, or
|
||
add an `until rclone lsf <artifact>; do sleep 60; done` preamble.
|
||
6. **Class scheduling**: `--class` drops jobs into
|
||
`gpuq/_class-pending/<class>/`; the single dispatcher (runs on the Spark,
|
||
`gpuq_dispatcher.sh`) moves each job to a concrete worker the moment one
|
||
is idle (empty pending + GPU util <15% + heartbeat fresher than 3 min).
|
||
Rosters: `gpuq/_classes/<class>.txt`, one worker id per line — edit with
|
||
`rclone`; adding/removing fleet nodes touches only these files. One
|
||
dispatcher only: it is the single writer that makes class queues
|
||
race-free. If it dies, class jobs simply wait; restart it anywhere.
|
||
7. Draining a node: stop submitting to its queues, wait for `--status` to
|
||
clear, `pkill -f gpuq_`, final `rclone copy` of its results dirs, then
|
||
destroy the instance. Pending jobs of a dead node survive in the bucket:
|
||
`rclone move jspace:jspace/gpuq/nodeX-gpuN/pending jspace:jspace/gpuq/nodeY-gpuM/pending`.
|