4.9 KiB
4.9 KiB
Where to find what
Two projects share this repo: the J-lens reproduction (does the 2026 workspace paper replicate on gemma-4-E2B?) and the workspace-looping investigation that grew out of it (retrofit recurrence onto the lens-found band; what does it actually buy?). The second is the active one.
The claims and their evidence
| you want | look in |
|---|---|
| Current claims, all numbers, figures | PAPER.md (source of truth; the .pdf snapshots lag it) |
| What worked / what failed / design rules / ops pitfalls | LESSONS.md — read before running anything on this hardware |
| Pre-registrations + scored outcomes (17 items, incl. refutations) | results-loop/PROTOCOL_UNIFIED.md — the methods backbone; every claim in PAPER.md §3 traces to an item here |
| Significance tests behind any claimed number | results-loop/STATS.md |
| Lab-notebook narrative of the looping investigation | WORKSPACE_LOOPING.md (superseded where it disagrees with PAPER.md) |
| Base-reproduction results (lens replication itself) | RESULTS.md, README.md |
| Per-model lens maps: workspace bands, KV-share boundaries, pinned revisions | results/REGIMES.json (canonical registry) + results/jbar*.pt (raw J̄) + results/exp4*.log (regime scans) |
Code (scripts/, jlens/)
jlens/core.py— the lens: model loading (JLENS_MODELenv), J̄ readouts.scripts/loop_common.py— everything band-looping:BandLooper(capture/re-run machinery, KV-cache-safe),generate_frozen_prompt(the ≥3.5× deploy path), and every adapter variant from the regime sweep (MergeAdapter★,AdaptiveMergeAdapter,RecurrentAdapter,ParcaeAdapter,NoisyMergeAdapter,TiedAlphaAdapter,PerDepthAdapter). Band viaJLENS_BANDenv (default E2B 14,30).- Trainers:
train_merge_code.py(MBPP; all regime flags live here),train_merge.py(GSM, old full-position regime — historic),train_merge_unified.py(multi-task, hardened protocol),train_merge_bw.py(Blocksworld),train_distill*.py(plan distillation). - Evals:
eval_loop_code.py(MBPP pass@1 vs k; per-item logs;--halt),eval_loop.py(GSM),eval_bw.py,eval_humaneval.py,eval_rust.py,eval_lcb.py,eval_mc_panel.py, plusprep_*.py(STaR labeling). - Figures:
fig_*.pyregenerate the canonical PNGs from the JSONs. - Infra:
gpuq_*.sh+GPUQ.md(bucket-backed GPU job queue),node_setup.sh(vast.ai bootstrap; pins model revisions — see LESSONS #12),vast-ai-notes.md.
Results directories — including the honest mess
results-loop/— the looping project's data: 84eval_*.json(tag suffixes:_s<seed>,_rec16/_parcae16recurrent arms,_pd4per-depth,_tatied-alpha,_rk16random-depth,_nsnoise-s₀,_h2048capacity,code2gsm_*transfer;per_itemonly in files from Jul 14 onward), adapter checkpoints (adapter_*.pt, e400 = the pre-committed eval checkpoint), canonical figures (fig_kcurves.pngdesign-space grid,fig_phase.pngtwo-dials diagram,fig_loop_vs_ff.pngrecurrence-vs-distill ladder,fig_placement/transfer/scale.png), andchain*.log— autonomous-session logs, archaeology only.results/— lens reproduction outputs + the cross-model registry (REGIMES.json).results-band-*/— one directory per entrance-placement arm of the placement sweep (L2–L24 entrances); summarized in PAPER fig_placement; kept for per-item audit.results-12b/,results-loop-12b/— 12B lens map and looping arms.results-tap23/,results-tap34/,results-kvtest/,results-combo/,results-panel/,results-distill-s7/— single-question side arms (exit-tap sweep, KV nulling check, combined arms, MC panel, distill seed).results-node*/,results-node2-final/— raw syncs from rented H100 nodes (500-item eval campaign).results-26b/— unclear provenance (Jul 13; layer indices ≤26 mean it is NOT the 26B MoE despite the name — possibly a misnamed early scan). Trust nothing here without re-derivation.paper-A/,paper-B/,paper-D/— abandoned paper-outline variants (one PLAN.md each); the live outline is PAPER.md itself.related_work/— the two anchor papers (McLeish 2511.07384, Lys 2602.14759), the workspace paper,relevant_to_us.mdnotes,bibliography.bib.
Conventions worth knowing
- Every eval prints a
k=0row first; it must equal the base model bit-exactly (0.488 on MBPP-250) — the sanity anchor that has caught two silent bugs (LESSONS #2, #12). - Difficulty labels (
easy/hard/drop) are STaR self-labels: direct-pass / CoT-only-pass / unreachable. "hard" = plan-dependent. - Checkpoints are pre-committed before evals (usually e400); post-hoc checkpoint shopping is flagged as exploratory wherever it happened.
- GPU jobs go through the gpuq queue (
gpuq_submit.sh <worker> <job.sh>), never bare nohup on the Spark; jobs are killed bypkill -fself-matches embarrassingly often (LESSONS #6).