Files
jspace/scripts/GPUQ.md
T

147 lines
6.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# gpuq — bucket-backed multi-node GPU job queue
A minimal scheduler for keeping N GPUs across M hosts busy when the hosts
cannot reach each other directly (rented vast.ai nodes) but all share one
S3 bucket (rclone remote `jspace`, Hetzner). Three shell scripts, no
daemon, no database, no cross-node networking.
```
scripts/gpuq_worker.sh one per GPU runs jobs from its own queue
scripts/gpuq_supervisor.sh one per node restarts workers, heartbeats
scripts/gpuq_submit.sh anywhere enqueue jobs / show status
```
## Design in one paragraph
Every GPU gets a private queue directory in the bucket:
`jspace:jspace/gpuq/<worker-id>/pending/`. A worker loop polls its own
directory (20 s), pulls job files, runs them oldest-first-by-name with
`CUDA_VISIBLE_DEVICES` pinned, and moves job + full log to `done/`.
There is **no work stealing and no job claiming** — each queue has exactly
one consumer, so the S3 lack of atomic move can never double-run a job.
Load balancing happens at submit time (you pick the worker); ordering
happens in the filename (`00_`, `01_`, …). The bucket is both transport
and durable state: a dead node's pending jobs stay visible and can be
moved to another worker with one `rclone move`.
## Bucket layout
```
gpuq/
_health/<node-id>.txt heartbeat, rewritten every 60 s
<node-id>-gpu<i>/
pending/<job>.sh waiting (consumed oldest-first by name)
done/<job>.sh finished job file
done/<job>.log its full stdout+stderr, ends with "rc=N"
```
Local mirror on the node: `~/gpuq/<worker-id>/` (pending copy, logs).
## Node setup (once per node, after git clone/pull of ~/jspace)
Requirements: `rclone` configured with the `jspace` remote (see
`.s3-credentials.txt` on the Spark — do NOT commit it), tmux session
`ssh_tmux`, GPUs visible to `nvidia-smi`.
Start ONLY the supervisor — it spawns and maintains all workers itself:
```bash
tmux new-window -t ssh_tmux -n gpuq_sup \
'bash ~/jspace/scripts/gpuq_supervisor.sh node1 5' # node-id, #GPUs
```
The supervisor, every 60 s:
1. respawns any missing `gpuq_worker.sh node1-gpu<i>` as tmux window
`gpuq<i>` (nohup fallback if tmux is unreachable);
2. publishes `gpuq/_health/node1.txt`: timestamp, uptime, per-GPU
util/mem, disk on `/` and `/dev/shm`, live worker count, local
pending count.
There is deliberately no supervisor-of-supervisors: a stale heartbeat
(> ~2 min) IS the signal that the node or supervisor died — check it
from outside.
## Job contract
A job is a plain bash file. The worker provides `CUDA_VISIBLE_DEVICES`;
everything else is the job's business. Template:
```bash
#!/bin/bash
cd ~/jspace/scripts
P=~/jspace/.venv/bin/python # /venv/main/bin/python on vast images
export HF_HUB_OFFLINE=1 # model already cached
LOOP_OUT=~/jspace/results-loop-12b $P train_merge_code.py --seed 7 --alpha 0.15
rclone copy ~/jspace/results-loop-12b jspace:jspace/results-12b/ \
--include "adapter_code_s7*" --include "*.json" -q
```
Rules of thumb:
- **Jobs sync their own results** to the bucket at the end (the queue only
ships the job's log automatically, not its artifacts).
- One GPU per job; never set CUDA_VISIBLE_DEVICES yourself.
- Idempotence is on the submitter: if you resubmit a job, it reruns.
- Exit code lands as the last line of the log (`rc=0`).
## Submitting and monitoring (from any machine with the remote)
```bash
# ordered pair on one GPU (train, then eval — order = filename sort):
scripts/gpuq_submit.sh node1-gpu0 00_train_a15.sh 01_eval_a15.sh
# fan five seeds across five GPUs:
for s in 0 1 2 3 4; do
scripts/gpuq_submit.sh node1-gpu$s seed$s.sh
done
scripts/gpuq_submit.sh --status # pending + last 20 finished
rclone cat jspace:jspace/gpuq/_health/node1.txt # node health
rclone cat jspace:jspace/gpuq/node1-gpu0/done/00_train_a15.log | tail -50
```
## Dependencies
- Same-GPU sequential dependency: filename prefixes on one worker
(`00_train.sh`, `01_eval.sh`). This covers nearly all our chains.
- Cross-GPU/cross-node dependencies: not supported by design. Either
submit the dependent job after seeing the first finish in `--status`,
or make the job itself poll the bucket for its input artifact before
starting (a `until rclone lsf ...; do sleep 60; done` preamble).
## Failure modes and recovery
| failure | effect | recovery |
|---|---|---|
| job crashes (rc≠0) | worker moves on to next job | read `done/<job>.log`; resubmit fixed job |
| worker dies | its queue stalls | supervisor respawns it ≤ 60 s |
| supervisor dies | workers keep running; heartbeat goes stale | restart supervisor window |
| node dies | heartbeat stale; pending jobs preserved in bucket | `rclone move gpuq/node1-gpu0/pending gpuq/node2-gpu3/pending` |
| duplicate submit | job runs twice (queues are dumb) | submitter's responsibility |
| partial upload read | impossible — S3 uploads are atomic (objects appear only complete) | — |
Local `~/gpuq/` is scratch; the bucket is the source of truth.
## Why not X (considered alternatives)
- **Slurm/K8s**: assume stable, mutually-reachable nodes; wrong shape for
ephemeral single-tenant rentals.
- **Ray / SkyPilot / dstack**: viable (dstack has native vast.ai support
and is the thing to try for the *next* provisioning burst), but they
own provisioning; gpuq feeds nodes that already exist.
- **iroh p2p binary**: would give push dispatch, live log streaming, and
direct node↔node transfer through NAT (QUIC hole-punching) — attractive
if the fleet becomes permanent. Rejected for now: jobs run 10 min2 h,
so poll latency is irrelevant, and durable queue state in S3 comes free
vs. a coordinator protocol we'd have to write and debug. For ad-hoc live
log streaming, n0's prebuilt `dumbpipe` works without writing code.
- **Marker-file watchers / pkill chains** (our previous approach): see
LESSONS.md — self-matching kill patterns and watcher pileups burned us
repeatedly. gpuq replaces markers with per-consumer queues and replaces
pkill with the supervisor owning worker lifecycle.
## Security notes
The bucket credentials on rented nodes are readable by the host operator:
keep the `jspace` remote scoped to this project's bucket, and rotate the
key after each rental burst (already flagged in vast-ai-notes.md).