8.2 KiB
gpuq — bucket-backed multi-node GPU job queue
A minimal scheduler for keeping N GPUs across M hosts busy when the hosts
cannot reach each other directly (rented vast.ai nodes) but all share one
S3 bucket (rclone remote jspace, Hetzner). Three shell scripts, no
daemon, no database, no cross-node networking.
scripts/gpuq_worker.sh one per GPU runs jobs from its own queue
scripts/gpuq_supervisor.sh one per node restarts workers, heartbeats
scripts/gpuq_submit.sh anywhere enqueue jobs / show status
Design in one paragraph
Every GPU gets a private queue directory in the bucket:
jspace:jspace/gpuq/<worker-id>/pending/. A worker loop polls its own
directory (20 s), pulls job files, runs them oldest-first-by-name with
CUDA_VISIBLE_DEVICES pinned, and moves job + full log to done/.
There is no work stealing and no job claiming — each queue has exactly
one consumer, so the S3 lack of atomic move can never double-run a job.
Load balancing happens at submit time (you pick the worker); ordering
happens in the filename (00_, 01_, …). The bucket is both transport
and durable state: a dead node's pending jobs stay visible and can be
moved to another worker with one rclone move.
Bucket layout
gpuq/
_health/<node-id>.txt heartbeat, rewritten every 60 s
<node-id>-gpu<i>/
pending/<job>.sh waiting (consumed oldest-first by name)
done/<job>.sh finished job file
done/<job>.log its full stdout+stderr, ends with "rc=N"
Local mirror on the node: ~/gpuq/<worker-id>/ (pending copy, logs).
Node setup (once per node, after git clone/pull of ~/jspace)
Requirements: rclone configured with the jspace remote (see
.s3-credentials.txt on the Spark — do NOT commit it), tmux session
ssh_tmux, GPUs visible to nvidia-smi.
Start ONLY the supervisor — it spawns and maintains all workers itself:
tmux new-window -t ssh_tmux -n gpuq_sup \
'bash ~/jspace/scripts/gpuq_supervisor.sh node1 5' # node-id, #GPUs
The supervisor, every 60 s:
- respawns any missing
gpuq_worker.sh node1-gpu<i>as tmux windowgpuq<i>(nohup fallback if tmux is unreachable); - publishes
gpuq/_health/node1.txt: timestamp, uptime, per-GPU util/mem, disk on/and/dev/shm, live worker count, local pending count.
There is deliberately no supervisor-of-supervisors: a stale heartbeat (> ~2 min) IS the signal that the node or supervisor died — check it from outside.
Job contract
A job is a plain bash file. The worker provides CUDA_VISIBLE_DEVICES;
everything else is the job's business. Template:
#!/bin/bash
cd ~/jspace/scripts
P=~/jspace/.venv/bin/python # /venv/main/bin/python on vast images
export HF_HUB_OFFLINE=1 # model already cached
LOOP_OUT=~/jspace/results-loop-12b $P train_merge_code.py --seed 7 --alpha 0.15
rclone copy ~/jspace/results-loop-12b jspace:jspace/results-12b/ \
--include "adapter_code_s7*" --include "*.json" -q
Rules of thumb:
- Jobs sync their own results to the bucket at the end (the queue only ships the job's log automatically, not its artifacts).
- One GPU per job; never set CUDA_VISIBLE_DEVICES yourself.
- Idempotence is on the submitter: if you resubmit a job, it reruns.
- Exit code lands as the last line of the log (
rc=0).
Submitting and monitoring (from any machine with the remote)
# ordered pair on one GPU (train, then eval — order = filename sort):
scripts/gpuq_submit.sh node1-gpu0 00_train_a15.sh 01_eval_a15.sh
# fan five seeds across five GPUs:
for s in 0 1 2 3 4; do
scripts/gpuq_submit.sh node1-gpu$s seed$s.sh
done
scripts/gpuq_submit.sh --status # pending + last 20 finished
rclone cat jspace:jspace/gpuq/_health/node1.txt # node health
rclone cat jspace:jspace/gpuq/node1-gpu0/done/00_train_a15.log | tail -50
Dependencies
- Same-GPU sequential dependency: filename prefixes on one worker
(
00_train.sh,01_eval.sh). This covers nearly all our chains. - Cross-GPU/cross-node dependencies: not supported by design. Either
submit the dependent job after seeing the first finish in
--status, or make the job itself poll the bucket for its input artifact before starting (auntil rclone lsf ...; do sleep 60; donepreamble).
Failure modes and recovery
| failure | effect | recovery |
|---|---|---|
| job crashes (rc≠0) | worker moves on to next job | read done/<job>.log; resubmit fixed job |
| worker dies | its queue stalls | supervisor respawns it ≤ 60 s |
| supervisor dies | workers keep running; heartbeat goes stale | restart supervisor window |
| node dies | heartbeat stale; pending jobs preserved in bucket | rclone move gpuq/node1-gpu0/pending gpuq/node2-gpu3/pending |
| duplicate submit | job runs twice (queues are dumb) | submitter's responsibility |
| partial upload read | impossible — S3 uploads are atomic (objects appear only complete) | — |
Local ~/gpuq/ is scratch; the bucket is the source of truth.
Why not X (considered alternatives)
- Slurm/K8s: assume stable, mutually-reachable nodes; wrong shape for ephemeral single-tenant rentals.
- Ray / SkyPilot / dstack: viable (dstack has native vast.ai support and is the thing to try for the next provisioning burst), but they own provisioning; gpuq feeds nodes that already exist.
- iroh p2p binary: would give push dispatch, live log streaming, and
direct node↔node transfer through NAT (QUIC hole-punching) — attractive
if the fleet becomes permanent. Rejected for now: jobs run 10 min–2 h,
so poll latency is irrelevant, and durable queue state in S3 comes free
vs. a coordinator protocol we'd have to write and debug. For ad-hoc live
log streaming, n0's prebuilt
dumbpipeworks without writing code. - Marker-file watchers / pkill chains (our previous approach): see LESSONS.md — self-matching kill patterns and watcher pileups burned us repeatedly. gpuq replaces markers with per-consumer queues and replaces pkill with the supervisor owning worker lifecycle.
Security notes
The bucket credentials on rented nodes are readable by the host operator:
keep the jspace remote scoped to this project's bucket, and rotate the
key after each rental burst (already flagged in vast-ai-notes.md).
As deployed (2026-07-14) — operator cheat sheet
Fleet:
| node | workers | supervisor lives in |
|---|---|---|
node3 (8×H100, vast) |
node3-gpu0 … node3-gpu7 |
tmux ssh_tmux:gpuq_sup |
node2 (2×H100, vast) |
node2-gpu0, node2-gpu1 |
tmux ssh_tmux:gpuq_sup (drain watcher armed on the Spark) |
spark (GB10) |
spark-gpu0 |
detached, tmux session gpuq_spark |
Daily usage (from the Spark or anywhere with the jspace remote):
bash scripts/gpuq_submit.sh node3-gpu4 myjob.sh # submit to a specific GPU
bash scripts/gpuq_submit.sh --status # pending + last 20 done
rclone cat jspace:jspace/gpuq/_health/node3.txt # heartbeat; stale >2 min = trouble
rclone cat jspace:jspace/gpuq/node3-gpu4/done/myjob.log | tail -30
Job conventions (learned the hard way — see LESSONS.md):
- First line of work:
cd <repo> && git pull origin main -q— jobs self-update; never rely on what a node cloned at boot (repo is public, no keys needed). HF_HUB_OFFLINE=1ONLY if every model/dataset the job touches is already cached on that node — it blocksdatasetsdownloads too (bit us twice).- Jobs end with their own
rclone copyof artifacts to the bucket; the queue ships only the log automatically. - Env the job must set itself: venv activation (
/venv/mainon vast,~/jspace/.venvon the Spark),JLENS_MODEL,JLENS_BAND,LOOP_OUT,HF_HOME. The worker provides exactly one thing:CUDA_VISIBLE_DEVICES. - Ordering on one GPU = filename sort (
00_train.sh,01_eval.sh). Cross-GPU deps: submit after the prerequisite shows in--status, or add anuntil rclone lsf <artifact>; do sleep 60; donepreamble. - Draining a node: stop submitting to its queues, wait for
--statusto clear,pkill -f gpuq_, finalrclone copyof its results dirs, then destroy the instance. Pending jobs of a dead node survive in the bucket:rclone move jspace:jspace/gpuq/nodeX-gpuN/pending jspace:jspace/gpuq/nodeY-gpuM/pending.