Files
jspace/vast-ai-notes.md
T
2026-07-14 01:14:57 +02:00

98 lines
5.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Vast.ai node runbook — lessons from the 2×H200 session (2026-07-14)
What actually works, and every trap we hit getting there. Cost of learning:
~80 idle GPU-minutes. Next setup should take <15 minutes.
## The recipe that works
```bash
# 1) image: Vast "PyTorch" template (torch in /venv/main, NOT system python)
source /venv/main/bin/activate
uv pip install -q "transformers==5.13.*" datasets accelerate hf_transfer
# 2) work in RAM — tmpfs beats the tiny container disk (check: df -h /dev/shm)
export HF_HOME=/dev/shm/hf
mkdir -p /dev/shm/jspace
# 3) code: rsync the LIVE working tree from the dev box — never a stale tarball
# (from dev box:) rsync -a -e "ssh -p PORT" jspace/scripts jspace/jlens root@IP:/dev/shm/jspace/
# 4) data/artifacts: rclone from the bucket (config: scp ~/.config/rclone/rclone.conf over)
rclone copy jspace:jspace/results-loop /dev/shm/jspace/results-loop/
# 5) models: hf_transfer or bust (see pitfall #1)
export HF_TOKEN=... HF_HUB_ENABLE_HF_TRANSFER=1
hf download google/gemma-4-12B-it # ~2.5 GB/s vs stalls without
# 6) AFTER downloads complete: force offline so nothing ever waits on network
export HF_HUB_OFFLINE=1 HF_DATASETS_OFFLINE=1
# (datasets caches rsync over from dev box: ~/.cache/huggingface/datasets/<name>)
# 7) EVERYTHING long-running goes in tmux windows — that is the deal
tmux new-window -t ssh_tmux -n myjob 'bash /dev/shm/run_myjob.sh'
# pipe stages through `tee` so panes show live output AND files exist for pollers
# 8) results survive teardown via a sync loop (tmpfs is volatile):
while true; do rclone copy /dev/shm/jspace/results-X jspace:jspace/results-X/; sleep 300; done
```
## Pitfalls, in the order they bit us
1. **Single-stream HF downloads stall dead** on some hosts (twice: at 9.6 GB
and 18 GB, zero error, zero timeout). `pip install hf_transfer` +
`HF_HUB_ENABLE_HF_TRANSFER=1` fixed it instantly (~2.5 GB/s). Always.
2. **Never kill a downloader mid-finalization.** Killing `hf download` while
it "verifies" left a snapshot missing `tokenizer.json` → prompts tokenized
to `<unk>`, generations decoded to `None`/empty, every labeling pass
silently produced 0.000 accuracy. Diagnosis that found it: `md5sum`/file
diff of the snapshot dir vs a known-good cache. A later `hf download`
re-run did NOT restore the missing file (it trusted the snapshot);
explicitly downloading the single file did:
`hf download <repo> tokenizer.json`.
3. **Stale `.locks` deadlock everything after any killed HF process.**
`rm -rf $HF_HOME/hub/.locks` before relaunching. Symptom: "Still waiting
to acquire lock… (elapsed: NNs)" forever.
4. **Two processes downloading the same model = lock contention.** Download
once, THEN start parallel jobs.
5. **`pkill -f <pattern>` kills your own command** if the pattern appears in
its command line (exit code 144, half your script never runs). Use
bracket-escaping: `pkill -f "[h]f download"`. Bit us three times, twice
locally, once remotely.
6. **Stale code on the node.** The bucket tarball was a day old; the node ran
pre-patch code (wrong adapter width, wrong band, missing env handling) and
produced invalid results that LOOKED like model problems. Rsync the live
tree; verify with a hash or a version marker if paranoid.
7. **nohup + ssh one-shots are fragile**; interrupted ssh calls orphan or
duplicate work, and you can't see what's happening. tmux windows in the
host's `ssh_tmux` session: survive disconnects, visible to the human
(`Ctrl-b <n>`), killable as a unit. Caveat: `Ctrl-C` in a pane kills the
whole window's process group, tail included.
8. **`.bashrc` on Vast auto-attaches tmux** — sourcing it in non-interactive
ssh fails ("duplicate session: ssh_tmux"). Set env inline instead.
9. **Model revisions/templates differ across a family.** gemma-4-12B-it uses
a channel-based chat template (`<|channel>thought`) and ends turns with
`<turn|>`, not E2B's `<end_of_turn>`. Generation-stop lists must be built
per-tokenizer (filter unk!), or numeric-answer extraction reads trailing
garbage. Probe ONE generation end-to-end (`convert_ids_to_tokens` on both
prompt and output) before batch-labeling anything on a new model.
10. **Probe-gate your launches.** The pattern that finally worked: a tmux
"fix" window that repairs → runs a single-generation probe → only on
probe-pass spawns the real jobs. (And make the pass-condition robust —
ours failed on a too-strict grep despite a correct generation.)
11. **`HF_HUB_OFFLINE=1` blocks dataset downloads too.** Ship the
`~/.cache/huggingface/datasets/<name>` dirs from the dev box (tiny), or
download datasets before going offline.
## Node facts worth remembering
- Vast = unprivileged Docker container on a shared host: GPUs exclusive,
CPU/RAM/network shared; the HOST OPERATOR can read everything on the box —
scope and rotate any credential that touches it (HF token, bucket keys).
- `vast-capabilities` prints the live manifest (ports, services, creds
presence). Read `/etc/vast-agents-guide.md` first.
- `/dev/shm` was 181 GB on a 1.4 TB-RAM box (container-capped) — still ample:
weights + caches + outputs all in RAM; instance restart loses it, hence the
rclone sync loop.
- 2×H200 NVL ran 12B labeling at ~98%/81% util with both branches parallel —
roughly the Spark's E2B speed on a 5× bigger model, per GPU.