vast.ai runbook + per-tokenizer eos handling in loop_common
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -240,9 +240,11 @@ class BandLooper:
|
|||||||
hook = self.tm.layers[self.l0].register_forward_pre_hook(
|
hook = self.tm.layers[self.l0].register_forward_pre_hook(
|
||||||
swap, with_kwargs=True)
|
swap, with_kwargs=True)
|
||||||
try:
|
try:
|
||||||
|
unk = tok.unk_token_id
|
||||||
eos = [t for t in (tok.eos_token_id,
|
eos = [t for t in (tok.eos_token_id,
|
||||||
tok.convert_tokens_to_ids("<end_of_turn>"))
|
tok.convert_tokens_to_ids("<end_of_turn>"),
|
||||||
if t is not None and t >= 0]
|
tok.convert_tokens_to_ids("<turn|>"))
|
||||||
|
if t is not None and t >= 0 and t != unk]
|
||||||
out = self.model.generate(
|
out = self.model.generate(
|
||||||
input_ids=input_ids, attention_mask=attention_mask,
|
input_ids=input_ids, attention_mask=attention_mask,
|
||||||
max_new_tokens=max_new_tokens, do_sample=False,
|
max_new_tokens=max_new_tokens, do_sample=False,
|
||||||
|
|||||||
@@ -0,0 +1,97 @@
|
|||||||
|
# Vast.ai node runbook — lessons from the 2×H200 session (2026-07-14)
|
||||||
|
|
||||||
|
What actually works, and every trap we hit getting there. Cost of learning:
|
||||||
|
~80 idle GPU-minutes. Next setup should take <15 minutes.
|
||||||
|
|
||||||
|
## The recipe that works
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1) image: Vast "PyTorch" template (torch in /venv/main, NOT system python)
|
||||||
|
source /venv/main/bin/activate
|
||||||
|
uv pip install -q "transformers==5.13.*" datasets accelerate hf_transfer
|
||||||
|
|
||||||
|
# 2) work in RAM — tmpfs beats the tiny container disk (check: df -h /dev/shm)
|
||||||
|
export HF_HOME=/dev/shm/hf
|
||||||
|
mkdir -p /dev/shm/jspace
|
||||||
|
|
||||||
|
# 3) code: rsync the LIVE working tree from the dev box — never a stale tarball
|
||||||
|
# (from dev box:) rsync -a -e "ssh -p PORT" jspace/scripts jspace/jlens root@IP:/dev/shm/jspace/
|
||||||
|
|
||||||
|
# 4) data/artifacts: rclone from the bucket (config: scp ~/.config/rclone/rclone.conf over)
|
||||||
|
rclone copy jspace:jspace/results-loop /dev/shm/jspace/results-loop/
|
||||||
|
|
||||||
|
# 5) models: hf_transfer or bust (see pitfall #1)
|
||||||
|
export HF_TOKEN=... HF_HUB_ENABLE_HF_TRANSFER=1
|
||||||
|
hf download google/gemma-4-12B-it # ~2.5 GB/s vs stalls without
|
||||||
|
|
||||||
|
# 6) AFTER downloads complete: force offline so nothing ever waits on network
|
||||||
|
export HF_HUB_OFFLINE=1 HF_DATASETS_OFFLINE=1
|
||||||
|
# (datasets caches rsync over from dev box: ~/.cache/huggingface/datasets/<name>)
|
||||||
|
|
||||||
|
# 7) EVERYTHING long-running goes in tmux windows — that is the deal
|
||||||
|
tmux new-window -t ssh_tmux -n myjob 'bash /dev/shm/run_myjob.sh'
|
||||||
|
# pipe stages through `tee` so panes show live output AND files exist for pollers
|
||||||
|
|
||||||
|
# 8) results survive teardown via a sync loop (tmpfs is volatile):
|
||||||
|
while true; do rclone copy /dev/shm/jspace/results-X jspace:jspace/results-X/; sleep 300; done
|
||||||
|
```
|
||||||
|
|
||||||
|
## Pitfalls, in the order they bit us
|
||||||
|
|
||||||
|
1. **Single-stream HF downloads stall dead** on some hosts (twice: at 9.6 GB
|
||||||
|
and 18 GB, zero error, zero timeout). `pip install hf_transfer` +
|
||||||
|
`HF_HUB_ENABLE_HF_TRANSFER=1` fixed it instantly (~2.5 GB/s). Always.
|
||||||
|
2. **Never kill a downloader mid-finalization.** Killing `hf download` while
|
||||||
|
it "verifies" left a snapshot missing `tokenizer.json` → prompts tokenized
|
||||||
|
to `<unk>`, generations decoded to `None`/empty, every labeling pass
|
||||||
|
silently produced 0.000 accuracy. Diagnosis that found it: `md5sum`/file
|
||||||
|
diff of the snapshot dir vs a known-good cache. A later `hf download`
|
||||||
|
re-run did NOT restore the missing file (it trusted the snapshot);
|
||||||
|
explicitly downloading the single file did:
|
||||||
|
`hf download <repo> tokenizer.json`.
|
||||||
|
3. **Stale `.locks` deadlock everything after any killed HF process.**
|
||||||
|
`rm -rf $HF_HOME/hub/.locks` before relaunching. Symptom: "Still waiting
|
||||||
|
to acquire lock… (elapsed: NNs)" forever.
|
||||||
|
4. **Two processes downloading the same model = lock contention.** Download
|
||||||
|
once, THEN start parallel jobs.
|
||||||
|
5. **`pkill -f <pattern>` kills your own command** if the pattern appears in
|
||||||
|
its command line (exit code 144, half your script never runs). Use
|
||||||
|
bracket-escaping: `pkill -f "[h]f download"`. Bit us three times, twice
|
||||||
|
locally, once remotely.
|
||||||
|
6. **Stale code on the node.** The bucket tarball was a day old; the node ran
|
||||||
|
pre-patch code (wrong adapter width, wrong band, missing env handling) and
|
||||||
|
produced invalid results that LOOKED like model problems. Rsync the live
|
||||||
|
tree; verify with a hash or a version marker if paranoid.
|
||||||
|
7. **nohup + ssh one-shots are fragile**; interrupted ssh calls orphan or
|
||||||
|
duplicate work, and you can't see what's happening. tmux windows in the
|
||||||
|
host's `ssh_tmux` session: survive disconnects, visible to the human
|
||||||
|
(`Ctrl-b <n>`), killable as a unit. Caveat: `Ctrl-C` in a pane kills the
|
||||||
|
whole window's process group, tail included.
|
||||||
|
8. **`.bashrc` on Vast auto-attaches tmux** — sourcing it in non-interactive
|
||||||
|
ssh fails ("duplicate session: ssh_tmux"). Set env inline instead.
|
||||||
|
9. **Model revisions/templates differ across a family.** gemma-4-12B-it uses
|
||||||
|
a channel-based chat template (`<|channel>thought`) and ends turns with
|
||||||
|
`<turn|>`, not E2B's `<end_of_turn>`. Generation-stop lists must be built
|
||||||
|
per-tokenizer (filter unk!), or numeric-answer extraction reads trailing
|
||||||
|
garbage. Probe ONE generation end-to-end (`convert_ids_to_tokens` on both
|
||||||
|
prompt and output) before batch-labeling anything on a new model.
|
||||||
|
10. **Probe-gate your launches.** The pattern that finally worked: a tmux
|
||||||
|
"fix" window that repairs → runs a single-generation probe → only on
|
||||||
|
probe-pass spawns the real jobs. (And make the pass-condition robust —
|
||||||
|
ours failed on a too-strict grep despite a correct generation.)
|
||||||
|
11. **`HF_HUB_OFFLINE=1` blocks dataset downloads too.** Ship the
|
||||||
|
`~/.cache/huggingface/datasets/<name>` dirs from the dev box (tiny), or
|
||||||
|
download datasets before going offline.
|
||||||
|
|
||||||
|
## Node facts worth remembering
|
||||||
|
|
||||||
|
- Vast = unprivileged Docker container on a shared host: GPUs exclusive,
|
||||||
|
CPU/RAM/network shared; the HOST OPERATOR can read everything on the box —
|
||||||
|
scope and rotate any credential that touches it (HF token, bucket keys).
|
||||||
|
- `vast-capabilities` prints the live manifest (ports, services, creds
|
||||||
|
presence). Read `/etc/vast-agents-guide.md` first.
|
||||||
|
- `/dev/shm` was 181 GB on a 1.4 TB-RAM box (container-capped) — still ample:
|
||||||
|
weights + caches + outputs all in RAM; instance restart loses it, hence the
|
||||||
|
rclone sync loop.
|
||||||
|
- 2×H200 NVL ran 12B labeling at ~98%/81% util with both branches parallel —
|
||||||
|
roughly the Spark's E2B speed on a 5× bigger model, per GPU.
|
||||||
Reference in New Issue
Block a user