Files

97 lines
7.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# What worked and what didn't — workspace looping on gemma-4-E2B
Lab-notebook distillation of the looping investigation (July 2026). Full
results and figures: [`WORKSPACE_LOOPING.md`](WORKSPACE_LOOPING.md); base
reproduction: [`RESULTS.md`](RESULTS.md). Everything on `google/gemma-4-E2B-it`
(frozen), DGX Spark GB10.
---
## Scientific: what WORKED
| finding | evidence | where |
|---|---|---|
| **Merge layer makes the band a stable recurrence** — anchor-dominant `(1−α)e + α·ŝ`, α=0.3; the whole failure of naive looping is that the band is not a self-map | untrained: answer preserved 11+ loops, converging (probe 2d) | `loop_common.py` MergeAdapter |
| **Training the merge (1.6M params, model frozen) turns hold into amplify** | J-lens concept sharpens 8× across loops (0.015→0.13; untrained flat) | `train_merge.py`, eval JSONs |
| **The loop is a true fixed-point iteration** — big first step, bit-exact convergence by k≈34; explains why accuracy & sharpening plateau there, k>4 is a no-op | cos(s_k, s_{k1}): 0.926 → 1.000 | `plot_loop_dynamics.py` |
| **Prompt-only looping ("latent planning") beats baseline on code** — first overall win: MBPP pass@1 52.0% vs 48.8%; plan-dependent problems 3.6%→46.4% (untrained control 1418%; p<1e-5) | 250 test items, executed | `eval_loop_code.py`, §5 |
| **Frozen-prompt KV-cache trick** — looped prompt states are constant across token steps (causality), so loop once, hook-swap the band input at prefill, generate cached. Bit-identical, ≥3.5× faster | equivalence test match=True | `generate_frozen_prompt` |
| **STaR self-labeling gives difficulty labels + in-distribution supervision for free** — direct pass vs CoT/plan pass, verified answers/tests | GSM8K 16%/53%, MBPP 52%/63% | `prep_star_data.py`, `prep_mbpp.py` |
| **Difficulty→depth curriculum is absorbed** — hard items (trained only at k≥2) fit best at depth | val CE gradient; hard-bucket peaks at k=4 | train logs |
| **Untrained cross-token carry is stable** — band state carried across token steps, one pass/token; coherent 60+ tokens, mild norm drift; mean/EMA seeding ≥ last-position | `probe_carry.py` output | §6 |
| **Trained sharpening transfers to the carry regime on short answers** — P('spider') 0.015→0.12 in ~2 token steps, amortized | carry probe w/ trained adapter | §6 |
## Scientific: what FAILED (and what each failure taught)
| failure | numbers | lesson |
|---|---|---|
| **Naive band loop** (L30 out → L14 in) | collapses in 1 iteration | the band is not a self-map; out-space ≠ in-space |
| **Additive anchor** `H0 + bandΔ(H)` | diverges, ‖ΔH‖/‖H‖≈2/step | forward displacement double-counts; need interpolation, not addition |
| **Frozen single tied layer** | stable but degenerate attractor; concept never sharpens | untrained layers tolerate looping, never benefit (matches Relaxed-Recursive) |
| **α ≥ 0.6 merge** | drifts to garbage | the anchor must dominate; the loop is a perturbation, not a replacement |
| **Full-position looping + teacher forcing (GSM8K)** | overall acc never beats k=0; easy items 97%→~35% in free-run while teacher-forced val looked great | exposure bias: generated tokens through a loop trained only on gold tokens. Loop placement must match deployment. Prompt-only looping fixed this *structurally* |
| **Trained (within-token) adapter in the carry regime, long outputs** | tag-soup / token garbage after ~1 sentence | adapters must be trained on the trajectory they produce; short-answer training does not license long free-running recurrence |
| **Uniform loop depth on mixed difficulty** | easy items pay ~613 points at k≥1 even trained | looping problems that need no extra compute only adds perturbation → per-prompt gate (probe on k=0 workspace state; STaR labels = free supervision) is the top v2 item |
## Operational pitfalls (each cost real time — don't re-hit)
1. **"Planning hurts" was a truncation artifact.** Plan pass at
`max_new_tokens=380` scored 17% vs 54% direct — every sampled output ran
out of budget mid-plan, before any code. At 700 tokens + terse-plan prompt,
planning adds ~11 points. *Check truncation before believing any
"CoT hurts small models" result.*
2. **`stop_strings="\n```"` matched the OPENING fence** (generations start
`\n```python`) → every generation halted at ~4 tokens → pass@1=0.000 at all
k, including k=0. The k=0 sanity anchor is what caught it. *Always have a
known-value row in every eval.*
3. **Full-vocab logits are the memory monster.** (B, T, 262k) tensors:
batch 32 eval + concurrent training = machine-wide OOM (kernel killed the
desktop session). Then `earlyoom` (installed on DGX Spark, trigger-happy
with unified memory) SIGTERM'd a modest eval *silently* — no traceback,
log just stops. Fixes: `logits_to_keep=1` during generation, batch ≤16,
**GPU jobs sequential**, and a watchdog that notifies when a process
disappears. Check `journalctl | grep earlyoom` for silent deaths.
4. **Re-calling decoder layers appends to the KV cache.** Capture-and-rerun
machinery must forward with `use_cache=False` and strip
`past_key_values` from captured kwargs, or re-runs see doubled keys.
Correctness anchor that caught it: manual band+suffix re-run must
reproduce the plain forward bit-exactly (0.0 max diff).
5. **Small hard pools overfit fast.** 38 hard items: val CE rises from ~step
300 of 600, depth-ordering inverts by step 500. Snapshot checkpoints and
evaluate the pre-overfit one (we used step 399); ~200 steps suffice.
6. **Background jobs must be `setsid`'d** or the harness/session restart
kills them mid-run. And `pkill -f <pattern>` will match your own launcher
shell if the pattern appears in its command line.
7. **Sustained training in a closed cabinet = thermal hard-freezes.** Four
crashes in one day (journal stops mid-line, no OOM, no shutdown trace,
37GB free at one death) on a DGX Spark that was stable all week under
light load. Pattern: only under hours of continuous GPU load; fixed by
opening the cabinet. Diagnose by exclusion: earlyoom quiet + journal
truncation + load correlation = thermal, not software. And do NOT
auto-restart training via @reboot cron on a thermally-suspect box — it
risks a crash loop with no human circuit breaker.
8. **Zero-init adapter output layer ⇒ zero grads upstream at step 0** — on
`mlp[0]` this is expected (LoRA-B-style), not a bug; check the output
layer's grad instead.
## The three design rules that emerged
1. **Anchor-dominant merge** makes any residual-stream recurrence well-posed;
train only the merge.
2. **Loop placement must match task structure**: prompt-only (static plan)
for generation tasks; full/carry (evolving state) for state-tracking —
and each must be *trained in the regime it deploys in*.
3. **Verify with the lens, gate with the labels**: the J-lens picks the band,
measures whether loops compute, and diagnoses failures; STaR difficulty
labels supervise both the curriculum and (next) the adaptive-depth gate.
## 12. Pin model revisions on fresh nodes
Upstream updated google/gemma-4-12B-it mid-project: the new chat template
adds a `<|channel>thought` scaffold, and greedy generation closes the empty
thought channel and stops — every generation decodes to "". Symptom:
0/500 pass rates with rc=0 (looks like a harness bug, is a silent model
swap). E2B was unaffected. Fix: `hf download --revision <hash>` + repoint
`refs/main` in the cache; node_setup.sh now pins both models (12B
0e2b1058…, E2B 9dbdf8a8…). Rule: any cross-node result assumes identical
model revisions — pin them, don't trust "main".