4.6 KiB
4.6 KiB
Paper A — Retrofitting Latent Planning onto Frozen LLMs
One-sentence claim. A 1.6M-parameter merge adapter at an interpretability-chosen layer boundary gives a frozen LLM a silent ("latent planning") mode that converts a large fraction of its plan-dependent failures into successes, at zero decode-time cost — and the gain is attributable to the recurrence, not the parameters.
Audience/venue. Main-track ML (ICLR/NeurIPS class); fallback: strong
workshop + arXiv. Base draft: ../PAPER.md (restructure, don't rewrite).
Claims and evidence manifest
| # | claim | evidence (files) | status |
|---|---|---|---|
| A1 | Hard-bucket transformation: 5.5%→37.4%±5.5, 5 seeds, full 500-item test | results-loop/eval_code_{code_s0_full,code_s1,code_s2,code_s3,code_s4}.json |
✅ |
| A2 | Overall never below baseline; +1.3±0.8, 5/5 seeds ≥ (sign p≈.03) | same as A1 | ✅ |
| A3 | Attribution: trained-noloop = untrained-loop = 17.9% ≪ trained loop | eval_code_{ff,untrained,trained}.json |
✅ |
| A4 | Pause-token control (token-compute ≠ loop) | eval_code_pause8.json |
⏳ today |
| A5 | Plan-distillation control (plan-info into weights ≠ loop) | eval_code_distill.json |
⏳ today |
| A6 | Explicit-plan reference: loop recovers ~43% of gap at 0 tokens | eval_plan_baseline.json |
✅ |
| A7 | Location matters (headline ablation only; full anatomy → Paper B): lens band 53.6/43.6 vs live wrong-bands ≤29/23.6; 12B clean version 32.8 vs 0.4–2.3 | results-band-*/eval*.json, results-node-final/results-12b/eval_12b_gsm_band10*.json |
✅ |
| A8 | Depth generalization k=3/6/8 (curriculum can't fake) | eval_code_code_kgen.json |
✅ |
| A9 | HumanEval transfer +8 overall; decomposition substrate +6.1 / training +1.3 | eval_humaneval_{trained,untrained}.json |
✅ |
| A10 | Bucket integrity: sampled relabeling 97% agreement; per-bucket n + Wilson CIs everywhere | mbpp_data.json (label_sampled), per_item logs |
✅ (recompute in stats pass) |
| A11 | Boundary section (math): no overall win at 2B or 12B; 4-arm grids; 12B attribution flip (loop>weights at scale); headroom law (gain tracks direct-vs-plan gap: 11→works, 8→works, 37→partial, 7→nothing) | eval_{untrained,uni,ff,carry,pausectl}.json, results-node-final/results-12b/eval_12b_gsm_*.json, eval_code_12b_* |
✅ |
| A12 | Deployment: frozen-prompt KV write-in ≈2.9× prefill/0 decode; gate preserves easy items | eval_gated.json, equivalence probe (§ scripts) |
✅ |
| A13 | Mechanism summary (2 paragraphs, cite Paper B): fixed point by k≈4; concept-probe sharpening 8×; KV-sharing caveat | loop_dynamics.png data, spider traces |
✅ |
Known weaknesses to state (not fix)
- One base-model family (gemma-4); E2B + 12B only. 12B MBPP fails as-trained
(α miscalibration shown by untrained 23.0% —
eval_code_12b_untrained.json); reported as scale-sensitivity of α, with GSM-12B carrying the scale story. - HumanEval transfer is substrate-dominant; framed as such (no overclaim).
- Lens-battery entropy metric uninformative (report honestly, keep concept-probe metric).
- MBPP contamination paragraph (both arms share it; hard bucket conservative).
Missing / nice-to-have
- ⏳ A4/A5 land today (night chain).
- Blocksworld + Rust transfer (today): include as "task-generality" table row whatever the outcome — flat/negative is fine, it feeds the headroom law.
- Optional strengthener (1 day, decide after draft): MBPP+ (EvalPlus strict tests) robustness row for seed-0.
- NOT required: second model family, rung-2 — future work section.
Outline (9 pages)
- Intro — token-space vs latent serial compute; the retrofit gap; contributions.
- Background — J-lens & workspace (cite Paper D), recurrent-depth line.
- Method — band location; anchor-dominant merge; STaR labeling; k-curriculum; prompt-only masking; frozen-prompt inference (fig: pipeline).
- Main results — A1/A2 (fig: 5-seed curves), A6, A9 (transfer), tables w/ CIs.
- Attribution — A3/A4/A5/A7/A8 grid (fig: attribution bars).
- Boundary — A11 (fig: 2×2 mechanism-placement grid, both scales).
- Mechanism summary — A13 (fig: dynamics + anchor cliff inset; details → Paper B).
- Deployment — A12; FLOPs framing.
- Related work, limitations, outlook (loopification-as-post-training thesis).
Work plan
- Stats pass over all per_item JSONs (Wilson, McNemar, sampled-label recompute) —
scripts/stats_pass.py, outputspaper-A/tables/. - Regenerate 5 figures into
paper-A/figures/. - Restructure PAPER.md →
paper-A/draft.mdper outline; PDF. - User adversarial-review round; revise.
- Owner of first draft: me, target: tomorrow evening.