Files
jspace/scripts/GPUQ.md
T

6.2 KiB
Raw Blame History

gpuq — bucket-backed multi-node GPU job queue

A minimal scheduler for keeping N GPUs across M hosts busy when the hosts cannot reach each other directly (rented vast.ai nodes) but all share one S3 bucket (rclone remote jspace, Hetzner). Three shell scripts, no daemon, no database, no cross-node networking.

scripts/gpuq_worker.sh      one per GPU     runs jobs from its own queue
scripts/gpuq_supervisor.sh  one per node    restarts workers, heartbeats
scripts/gpuq_submit.sh      anywhere        enqueue jobs / show status

Design in one paragraph

Every GPU gets a private queue directory in the bucket: jspace:jspace/gpuq/<worker-id>/pending/. A worker loop polls its own directory (20 s), pulls job files, runs them oldest-first-by-name with CUDA_VISIBLE_DEVICES pinned, and moves job + full log to done/. There is no work stealing and no job claiming — each queue has exactly one consumer, so the S3 lack of atomic move can never double-run a job. Load balancing happens at submit time (you pick the worker); ordering happens in the filename (00_, 01_, …). The bucket is both transport and durable state: a dead node's pending jobs stay visible and can be moved to another worker with one rclone move.

Bucket layout

gpuq/
  _health/<node-id>.txt            heartbeat, rewritten every 60 s
  <node-id>-gpu<i>/
    pending/<job>.sh               waiting (consumed oldest-first by name)
    done/<job>.sh                  finished job file
    done/<job>.log                 its full stdout+stderr, ends with "rc=N"

Local mirror on the node: ~/gpuq/<worker-id>/ (pending copy, logs).

Node setup (once per node, after git clone/pull of ~/jspace)

Requirements: rclone configured with the jspace remote (see .s3-credentials.txt on the Spark — do NOT commit it), tmux session ssh_tmux, GPUs visible to nvidia-smi.

Start ONLY the supervisor — it spawns and maintains all workers itself:

tmux new-window -t ssh_tmux -n gpuq_sup \
  'bash ~/jspace/scripts/gpuq_supervisor.sh node1 5'   # node-id, #GPUs

The supervisor, every 60 s:

  1. respawns any missing gpuq_worker.sh node1-gpu<i> as tmux window gpuq<i> (nohup fallback if tmux is unreachable);
  2. publishes gpuq/_health/node1.txt: timestamp, uptime, per-GPU util/mem, disk on / and /dev/shm, live worker count, local pending count.

There is deliberately no supervisor-of-supervisors: a stale heartbeat (> ~2 min) IS the signal that the node or supervisor died — check it from outside.

Job contract

A job is a plain bash file. The worker provides CUDA_VISIBLE_DEVICES; everything else is the job's business. Template:

#!/bin/bash
cd ~/jspace/scripts
P=~/jspace/.venv/bin/python            # /venv/main/bin/python on vast images
export HF_HUB_OFFLINE=1                # model already cached
LOOP_OUT=~/jspace/results-loop-12b $P train_merge_code.py --seed 7 --alpha 0.15
rclone copy ~/jspace/results-loop-12b jspace:jspace/results-12b/ \
  --include "adapter_code_s7*" --include "*.json" -q

Rules of thumb:

  • Jobs sync their own results to the bucket at the end (the queue only ships the job's log automatically, not its artifacts).
  • One GPU per job; never set CUDA_VISIBLE_DEVICES yourself.
  • Idempotence is on the submitter: if you resubmit a job, it reruns.
  • Exit code lands as the last line of the log (rc=0).

Submitting and monitoring (from any machine with the remote)

# ordered pair on one GPU (train, then eval — order = filename sort):
scripts/gpuq_submit.sh node1-gpu0 00_train_a15.sh 01_eval_a15.sh

# fan five seeds across five GPUs:
for s in 0 1 2 3 4; do
  scripts/gpuq_submit.sh node1-gpu$s seed$s.sh
done

scripts/gpuq_submit.sh --status          # pending + last 20 finished
rclone cat jspace:jspace/gpuq/_health/node1.txt     # node health
rclone cat jspace:jspace/gpuq/node1-gpu0/done/00_train_a15.log | tail -50

Dependencies

  • Same-GPU sequential dependency: filename prefixes on one worker (00_train.sh, 01_eval.sh). This covers nearly all our chains.
  • Cross-GPU/cross-node dependencies: not supported by design. Either submit the dependent job after seeing the first finish in --status, or make the job itself poll the bucket for its input artifact before starting (a until rclone lsf ...; do sleep 60; done preamble).

Failure modes and recovery

failure effect recovery
job crashes (rc≠0) worker moves on to next job read done/<job>.log; resubmit fixed job
worker dies its queue stalls supervisor respawns it ≤ 60 s
supervisor dies workers keep running; heartbeat goes stale restart supervisor window
node dies heartbeat stale; pending jobs preserved in bucket rclone move gpuq/node1-gpu0/pending gpuq/node2-gpu3/pending
duplicate submit job runs twice (queues are dumb) submitter's responsibility
partial upload read impossible — S3 uploads are atomic (objects appear only complete)

Local ~/gpuq/ is scratch; the bucket is the source of truth.

Why not X (considered alternatives)

  • Slurm/K8s: assume stable, mutually-reachable nodes; wrong shape for ephemeral single-tenant rentals.
  • Ray / SkyPilot / dstack: viable (dstack has native vast.ai support and is the thing to try for the next provisioning burst), but they own provisioning; gpuq feeds nodes that already exist.
  • iroh p2p binary: would give push dispatch, live log streaming, and direct node↔node transfer through NAT (QUIC hole-punching) — attractive if the fleet becomes permanent. Rejected for now: jobs run 10 min2 h, so poll latency is irrelevant, and durable queue state in S3 comes free vs. a coordinator protocol we'd have to write and debug. For ad-hoc live log streaming, n0's prebuilt dumbpipe works without writing code.
  • Marker-file watchers / pkill chains (our previous approach): see LESSONS.md — self-matching kill patterns and watcher pileups burned us repeatedly. gpuq replaces markers with per-consumer queues and replaces pkill with the supervisor owning worker lifecycle.

Security notes

The bucket credentials on rented nodes are readable by the host operator: keep the jspace remote scoped to this project's bucket, and rotate the key after each rental burst (already flagged in vast-ai-notes.md).