354 lines
20 KiB
Markdown
354 lines
20 KiB
Markdown
# Latent Planning by Workspace Recurrence: an Interpretability-Placed Implant, and What It Actually Buys
|
||
|
||
*Final-data draft, 2026-07-14. Base models: google/gemma-4-E2B-it and
|
||
gemma-4-12B-it, both frozen. Hardware: DGX Spark + rented 2×/8×H100 nodes.
|
||
Code, per-item logs, and pre-registrations: `~/jspace` (git). Statistics:
|
||
`results-loop/STATS.md`.*
|
||
|
||
## Abstract
|
||
|
||
Interpretability work with an averaged-Jacobian lens ("J-lens") partitions a
|
||
pretrained language model's depth into regimes, including a mid-depth
|
||
*workspace* band that holds verbalizable, unspoken intermediate content. We
|
||
retrofit recurrence onto this band in a **frozen** model: a 1.6M-parameter
|
||
anchor-dominant merge adapter (0.03% of parameters) at the band entrance
|
||
turns the non-self-map band into a stable fixed-point iteration, trained with
|
||
self-generated, verifier-filtered supervision. Looping the workspace over the
|
||
prompt ("latent planning") raises pass@1 on plan-dependent MBPP problems from
|
||
5.5% to 43.6% (seed mean 37.5±5.5), with zero visible tokens and zero
|
||
additional decode cost. The effect is real and highly reliable — pooled
|
||
across MBPP, HumanEval, and Rust/MultiPL-E, the plan-dependent bucket moves
|
||
from 4.2% to 35.6% (McNemar p≈1.5e-10) — and **placement is a law, not a
|
||
convenience**: the gain appears only when the loop enters at the
|
||
lens-identified boundary (L14), collapsing at L13 and below and structurally
|
||
nulling above.
|
||
|
||
But a complete attribution program deflates the mechanism's mystique.
|
||
(i) **The loop's content is amortizable**: distilling the model's own
|
||
explicit plans into the same-size adapter — no recurrence at inference —
|
||
matches or exceeds the loop on the same bucket (mean over 8 runs 45.7±4.6 vs
|
||
37.5±5.5, paired difference n.s.), and the two do not stack; running the
|
||
loop on top of the distilled adapter *degrades* it. (ii) **Width rivals
|
||
depth**: 16 trained pause registers reach 36.4% on the same bucket.
|
||
(iii) **Compute-matched token baselines are uncomfortable**: best-of-3
|
||
sampling beats every latent arm on overall accuracy (57.2% vs ≤55.2%), and a
|
||
50-token visible plan matches the loop on the hard bucket (40.0%). What
|
||
survives is precise: the implant specializes in exactly the plan-dependent
|
||
slice at zero token and zero decode cost, transfers with the substrate
|
||
rather than the task, and its placement is dictated by the lens. At 12B a
|
||
constant merge coefficient destroys the substrate; making the coefficient
|
||
state-dependent (a 3.8K-parameter gate) restores it on MBPP
|
||
(hard 11.4%→27.3% with overall preserved) but not on Blocksworld or GSM8K —
|
||
the stability dial that unifies this work with McLeish et al. (2511.07384)
|
||
and Lys et al. (2602.14759) is task- and scale-dependent.
|
||
|
||
## 1. What this paper claims
|
||
|
||
1. **A placement law.** The retrofit works if and only if the recurrence
|
||
enters at the lens boundary. Entrances at L9–L13 (same adapter, data,
|
||
curriculum) destroy overall accuracy (14–34% vs 52%) while recovering at
|
||
most half the hard-bucket gain; entrance at L14 preserves overall and
|
||
maximizes the gain (fig_placement). Entrances at L17/L24 are *structurally
|
||
null* in this architecture: KV-sharing makes layers ≥15 reuse keys/values
|
||
computed at ≤14, so k>0 is bit-identical to k=0 — a hazard for any
|
||
retrofit method that skips the mechanistic check. Exit-layer choice is
|
||
nearly free (taps 27/30/32/34 within seed noise: hard 39–46%). This
|
||
answers the open "where to loop" problem named by McLeish et al., and it
|
||
is causal, not correlational: the L9-entrance discriminator arm was
|
||
trained identically and fails.
|
||
|
||
2. **A verified, statistically solid capability gain on a narrow slice.**
|
||
Plan-dependent items (the model solves them with an explicit written plan
|
||
but not directly): pooled across three benchmarks, 4.2%→35.6%,
|
||
p≈1.5e-10. Overall accuracy is statistically unchanged on MBPP
|
||
(p=0.34) and improved on HumanEval transfer (58.5%→66.5%, p=0.011).
|
||
|
||
3. **A deflationary mechanism finding.** The trained loop converges to a
|
||
fixed point by k≈3–4 and behaves as *amortized plan content*, not
|
||
iterative computation: plan-distillation into the identical architecture
|
||
without recurrence matches it; stacking buys nothing (loop-training a
|
||
distill-warmed adapter: 34.5%, below distill alone; running the distilled
|
||
adapter in loop mode: drops to 20.0%); deeper k at inference is flat
|
||
(k=8: 40.0%). The recurrence is a *training-time scaffold* that lets the
|
||
adapter find plan-shaped content — content that can equally be put there
|
||
by distillation if plans are available.
|
||
|
||
4. **A width-vs-depth law.** Trained pause registers (width) capture most of
|
||
the plan effect on code; recurrence (depth) is needed only where a state
|
||
must *evolve* — on GSM8K generation-side carry beats registers, and on
|
||
Blocksworld (pure planning, no world knowledge) the loop lifts hard-split
|
||
plans 0%→43% at 2B where everything else fails. Plans are wide; execution
|
||
is deep.
|
||
|
||
5. **Honest economics.** The implant's costs: ≈2.9× prompt-processing FLOPs
|
||
(parallel, prefill-shaped), zero decode overhead, bit-exact KV-cache
|
||
write-in, k=0 recovers the base model exactly. Its competition at matched
|
||
FLOPs: best-of-3 sampling wins overall accuracy outright (57.2%); a
|
||
50-token visible plan ties the hard bucket. The value proposition is
|
||
*only*: no visible tokens, no decode latency, and the hard-slice
|
||
specialization (distill's 46% > budget-CoT's 40% > best-of-3's 33%).
|
||
|
||
6. **Scale transfers only with a state-dependent stability dial.** At 12B the
|
||
2B-tuned constant α=0.3 collapses overall accuracy (72.6%→43.0%); the
|
||
damage is present *before* adapter training (untrained-loop arm) and is
|
||
not fixed by retuning α or LR. A per-position learned coefficient
|
||
α=σ(w·[e;ŝ]+b) restores MBPP (overall 69.4%, hard 11.4%→27.3%) — but
|
||
fails to rescue Blocksworld-12B and yields only a marginal GSM8K-12B
|
||
overall gain (35.9%→36.7% at k=1), the project's only overall 12B win.
|
||
|
||
## 2. Method
|
||
|
||
**Locating the band.** The lens reads residual state h at layer ℓ through the
|
||
averaged Jacobian J̄_ℓ = E[∂h_final/∂h_ℓ] and the unembedding; depth regimes
|
||
follow from what the readout tracks. On gemma-4-E2B: workspace ≈ L14–30 of
|
||
35; on 12B: L36–45 of 48.
|
||
|
||
**Making the band a self-map.** Feeding L30's output back to L14 collapses
|
||
(out-space ≠ in-space). With e = L13's output (fixed anchor) and s the
|
||
fed-back, norm-matched band output:
|
||
|
||
L14-in = (1−α)·e + α·ŝ + MLP([e ; ŝ]), ŝ = s·|e|/|s|
|
||
|
||
α=0.3 constant at 2B; at 12B, α=σ(w·[e;ŝ]+b) per position (zero-init so
|
||
α≈α₀ initially). MLP output zero-init: the untrained adapter is exactly the
|
||
hand merge — stable, answer-preserving, content-holding.
|
||
|
||
**Training.** STaR-style self-labeling: items the frozen model solves only
|
||
with an explicit plan/CoT are "hard", direct solves "easy", neither "drop".
|
||
Cross-entropy on answer/code tokens of the direct prompt, model's own
|
||
verified outputs as targets. Difficulty→depth curriculum (easy k=1, mixed
|
||
k=2, hard k=2–4). The loop applies to the **prompt span only**; generated
|
||
tokens run the plain path but attend to looped prompt states. Variants
|
||
trained the same way: **pause-N** (N trained register tokens appended to the
|
||
prompt, no recurrence), **plan-distill** (KL from the model's own
|
||
plan-in-context distribution into the FF adapter), **rung-2** (warm-started
|
||
adapter + entrance-faded LoRA rank 8 on the band's first layers, loop-only
|
||
via a global toggle), and **stack** arms (distill-warm + loop training;
|
||
distilled adapter evaluated in loop mode).
|
||
|
||
**Inference.** Looped prompt states are causally independent of generated
|
||
tokens: computed once at prefill, written into the KV cache by a hooked
|
||
forward pass, generation native. Verified bit-identical to the slow path.
|
||
Cost at k=4: ≈2.9× prefill FLOPs, **zero** decode overhead.
|
||
|
||
## 3. Results
|
||
|
||
Statistics throughout: Wilson 95% CIs; paired comparisons by exact McNemar;
|
||
all headline arms evaluated on the full 500-item MBPP test split (hard
|
||
bucket n=55), HumanEval n=164 (hard n=38), Rust/MultiPL-E n=154 (hard n=25),
|
||
execution-verified. Label robustness: redefining "hard" as
|
||
labeled-hard ∧ k=0-fails-in-all-five-seeds (52/55 items) moves headline
|
||
numbers <2 points.
|
||
|
||
### 3.1 The placement law
|
||
|
||

|
||
|
||
Entrance-layer sweep with everything else fixed. L14 (lens boundary):
|
||
hard 43.6%, overall 53.6%. L13: hard 17.9%, overall 34.4%. L9–L12: overall
|
||
14.0–30.8% (substrate destroyed). L17/L24 entrances: k>0 ≡ k=0 (KV sharing;
|
||
verified bit-identical) — the 12B model has no shared-KV layers, making it
|
||
the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34):
|
||
hard 39.3–46.4%, within seed spread. The lens boundary is necessary; the
|
||
exit is a free parameter. (The L23-exit arm died in training and was not
|
||
rerun; the four completed exits bracket it.)
|
||
|
||
### 3.2 The attribution ladder
|
||
|
||

|
||
|
||
MBPP hard bucket (plan-dependent, n=55 unless noted):
|
||
|
||
| arm | hard pass@1 | overall |
|
||
|---|---|---|
|
||
| base (k=0, bit-exact) | 5.5% | 51.8% |
|
||
| untrained loop (α-merge only, n=28) | 17.9% | ~52% |
|
||
| trained FF, no recurrence (n=28) | 17.9% | ~52% |
|
||
| pause-16 registers (width) | 36.4% | 55.2% |
|
||
| **trained loop k=4** (seed mean, 5 seeds) | **37.5±5.5** (best 43.6) | 53.6% |
|
||
| rung-2: + entrance-faded band LoRA (n=28) | 42.9/46.4 (2 seeds) | 51.2/52.4 |
|
||
| **plan-distilled FF** (mean, 8 runs) | **45.7±4.6** (best 49.1) | 55.5% |
|
||
| budget-CoT (50 visible tokens) | 40.0% | 53.8% |
|
||
| best-of-3 sampling (≈matched FLOPs) | 32.7% | **57.2%** |
|
||
| explicit plan in context (ceiling) | 94.5% | 59.0% |
|
||
|
||
Significance structure (McNemar, `STATS.md`): loop vs base on hard,
|
||
p=5.7e-6; every latent-arm-vs-latent-arm difference (loop vs distill, distill
|
||
vs stack) is **not significant** at n=55; loop vs base *overall* is not
|
||
significant on MBPP (p=0.34). The ladder's shape is reliable; its fine
|
||
ordering is not.
|
||
|
||
### 3.3 The decisive tests: nothing stacks
|
||
|
||
If the loop performed genuine iterative computation, plan-distilled content
|
||
plus recurrence should compound. It does not:
|
||
|
||
- **Distill-warm + loop training**: hard 34.5% — below distill alone.
|
||
- **Distilled adapter run in loop mode**: hard 20.0%, overall 45.8% —
|
||
looping *degrades* the distilled weights.
|
||
- **Pause-16 + distill**: hard 30.9% — no width stacking either.
|
||
- **Inference depth beyond convergence**: k=8 hard 40.0% ≈ k=4 (fixed point,
|
||
cos(sₖ,sₖ₋₁)=1.000 by k≈3–4).
|
||
|
||
Reading: the recurrence is a **training-time scaffold**. The curriculum
|
||
forces hard-item loss to be reducible only through the loop, and what the
|
||
adapter learns to inject is plan-shaped content — the same content
|
||
distillation installs directly when explicit plans are available. The loop's
|
||
distinctive value is that it finds this content *without* plan supervision
|
||
(STaR labels only say which items needed plans, not what the plans were).
|
||
|
||
### 3.4 Compute-matched honesty
|
||
|
||
At approximately matched FLOPs, token-space baselines are strong: best-of-3
|
||
sampling wins overall accuracy against every latent arm (57.2%,
|
||
CI [52.8, 61.5], vs loop 53.6 [49.2, 57.9] — point estimate higher, CIs
|
||
overlap) by preserving easy items perfectly while sampling rescues some hard
|
||
ones. A 50-token visible plan ties the loop's hard bucket. The latent
|
||
implant's surviving advantages are qualitative: zero visible tokens (silent),
|
||
zero decode overhead (prefill-parallel; sampling and CoT pay serially at
|
||
bandwidth-bound decode), and the hard-slice crown under distillation (46% vs
|
||
40% budget-CoT vs 33% best-of-3). For deployment this means: the implant is
|
||
a *latency/token-budget* technology with a side specialization in
|
||
plan-dependent items — not an accuracy technology.
|
||
|
||
### 3.5 Width vs depth, and the task boundary
|
||
|
||
Pause registers (width) reach 36.4% (16 registers; 8: 30.9%, 32: 34.5% — flat
|
||
in N) on MBPP hard: static plan content fits in registers. GSM8K inverts the
|
||
prompt-side result entirely (no variant beats the weights control
|
||
prompt-side), but generation-side *carry* — recurrence across token steps —
|
||
doubles the pause control on hard items: arithmetic's serial state evolves
|
||
during the answer. Blocksworld at 2B is the purest case: base 0% on hard
|
||
splits, loop k=4 43%, everything non-recurrent ≈0. The law: **plans are
|
||
wide; execution is deep.** Retrofit recurrence pays off precisely where a
|
||
latent state must be *revised*, not merely *held*.
|
||
|
||
### 3.6 Scale: the stability dial
|
||
|
||

|
||
|
||
At 12B (no shared KV — unconfounded), constant α=0.3: overall collapses
|
||
72.6%→43.0% at k=4 while hard limps to 11.4%. The untrained-loop arm shows
|
||
the damage precedes adapter training; α=0.15 and LR retuning do not fix it
|
||
(47.6/52.6% overall). The state-dependent coefficient does, on MBPP:
|
||
overall 69.4% (base 72.4%), hard 11.4%→27.3%. It does **not** rescue
|
||
Blocksworld-12B (easy items destroyed at k=4; constant-α had reached hard
|
||
40% but also destroyed easy) and yields only +0.8 points overall on
|
||
GSM8K-12B (35.9→36.7 at k=1, hard 1.6→10.6) — the sole overall-accuracy win
|
||
of the program, and a marginal one. Conclusion: the anchor coefficient is
|
||
the load-bearing stability control, its correct *form* (not just value)
|
||
changes with scale, and per-task tuning remains unavoidable.
|
||
|
||
### 3.7 Transfer: substrate, not task
|
||
|
||

|
||
|
||
MBPP-trained implants applied unchanged: **HumanEval** overall 58.5%→66.5%
|
||
(loop k=4, p=0.011; hard 0→31.6%); notably the *untrained* merge already
|
||
reaches 64.6% and the transferred pause adapter 66.5% (hard 38.9%) — the
|
||
transfer is substrate-shaped (a generically useful perturbation+content
|
||
mode), not task-memorized. **Rust/MultiPL-E** (Python-trained, different
|
||
language, compile-run-verified): hard 8.0%→24.0% (p=0.125 at n=25 —
|
||
directionally consistent, underpowered). **Blocksworld** MBPP-transfer:
|
||
hard 0→14.3% (task-trained: 43%). Content transfers where the substrate's
|
||
plan-representation overlaps; task-specific training still dominates.
|
||
|
||
### 3.8 Mechanism, verification, deployment
|
||
|
||
The trained loop takes a large first step (cos(s₁,s₀)=0.926 vs 0.977
|
||
untrained) and converges bit-exactly by k≈3–4; accuracy and lens-sharpening
|
||
plateau there. P(latent concept) under the J-lens at the band exit rises
|
||
0.015→0.13 across iterations (~8× the untrained hold) — the lens that placed
|
||
the implant also renders its silent content inspectable. The STaR labels
|
||
train a free difficulty gate (route predicted-hard to k=4, else k=0);
|
||
gate quality (19% precision at 64% recall) is the current ceiling on
|
||
removing the easy-item perturbation tax. k=0 is the exact base model by
|
||
construction — the implant is removable at token granularity.
|
||
|
||
**General-capability panel** (ARC-Challenge, WinoGrande, HellaSwag, MMLU;
|
||
length-normalized MC scoring with the loop applied to the context span) is
|
||
running on the Spark; results will quantify what k>0 does to off-task
|
||
abilities. [PENDING — fill on completion.]
|
||
|
||
### 3.9 Negative results with content
|
||
|
||
Mixed-task (code+math) training regressed both tasks at equal validation CE
|
||
— CE parity does not predict generation parity, and validation-CE checkpoint
|
||
selection fails likewise (fixed-step pre-commitment used instead; no
|
||
checkpoint was selected on test or generation results). GSM8K distillation
|
||
collapsed to empty outputs twice (E2B first attempt, 12B) on 3-token targets
|
||
under KL-dominant loss; a CE-dominant retry at E2B trained but reached only
|
||
hard 4.7%. Plan-distillation on GSM8K underperforms its MBPP twin even when
|
||
training succeeds: consistent with §3.5, there is little static plan content
|
||
for math to amortize.
|
||
|
||
## 4. Related work
|
||
|
||
**McLeish et al. (arXiv:2511.07384)** retrofit depth-recurrence via layer
|
||
surgery + ~50B-token continued pretraining of all parameters; they name
|
||
layer choice as an open problem — §3.1 is a causal answer. Their surgery
|
||
needs a healing phase; our k=0 is exactly the base model. **Lys et al.
|
||
(arXiv:2602.14759)** loop frozen models training-free; their finding that
|
||
naive looping degrades while interpolation with the un-looped state rescues
|
||
it is independent convergent evidence for anchor-dominance, and their
|
||
setting is the untrained cell of our ladder (17.9%).
|
||
|
||
**One mechanism, three regimes.** All three works mix the fed-back state
|
||
with an anchor from the un-looped computation. Lys et al.'s moving average
|
||
η·h⁽⁰⁾+(1−η)·h⁽ᵗ⁾ is an untrained anchor coefficient; our
|
||
(1−α)e + α·ŝ + MLP is its trained analogue; McLeish et al.'s input injection
|
||
is the fully-learned limit. The 12B episode closes the loop on this
|
||
unification: the coefficient is the stability dial, naive looping is its
|
||
α→1 collapse limit, and our scale failure + state-dependent fix show the
|
||
dial must itself become a function of the state as models grow. Our stacking
|
||
results add a caution for the whole family: if retrofitted recurrence
|
||
content is amortizable (§3.3), some of the family's gains may be
|
||
reproducible by distillation without inference-time recurrence — a control
|
||
neither bracket paper runs.
|
||
|
||
Earlier lineage: Universal Transformers; DEQ; Huginn (2502.05171);
|
||
Mixture-of-Recursions (2507.10524); Relaxed Recursive Transformers
|
||
(2410.20672); Coconut; pause tokens (Goyal et al.) — whose trained variant
|
||
proved a genuine rival, not a strawman (§3.2, §3.5).
|
||
|
||
What remains distinct here: interpretability-derived placement with causal
|
||
validation; a fully frozen base with bit-exact k=0 and zero-decode-cost KV
|
||
write-in; the complete attribution ladder including compute-matched
|
||
token-space baselines and stacking tests; the width/depth task law; and the
|
||
amortizability finding itself.
|
||
|
||
## 5. Limitations
|
||
|
||
One model family (gemma-4), two scales, three task families. Hard buckets
|
||
are small (n=55/38/25); within-ladder orderings are not individually
|
||
significant, and only the pooled hard effect and the HumanEval overall gain
|
||
survive multiple-comparison scrutiny. Bucket membership derives from greedy
|
||
labeling runs (consensus-k0 robustness check moves numbers <2 points, but
|
||
both checks share the base model). Best-of-3/budget-CoT lack per-item logs
|
||
(no paired tests against them). The L23 exit arm and a third architecture
|
||
family were not run; LiveCodeBench (contamination-safe) was not run; rung-2
|
||
was not run at 12B. The easy-item perturbation tax persists wherever the
|
||
gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both
|
||
arms share contamination, and memorized items land in the easy bucket, but
|
||
bucket composition is contamination-sensitive. The capability panel
|
||
(§3.8) is pending; until it lands, off-task effects of k>0 are unmeasured.
|
||
The Blocksworld-12B and GSM8K-12B failures mean the adaptive-α fix is
|
||
demonstrated on one task at one scale, not established as a general recipe.
|
||
|
||
## 6. Conclusion
|
||
|
||
The experiment this program set out to run — *can an interpretability lens
|
||
tell you where to install recurrence in a frozen model, and does it work?* —
|
||
has a clean answer: yes, and the placement is causally load-bearing. The
|
||
more interesting answer is what the recurrence turned out to be: not a
|
||
reasoning engine, but a remarkably cheap way to make a frozen model amortize
|
||
its own planning into 0.03% of extra parameters, with a training-time loop
|
||
as scaffold and an inference-time loop that is optional once the content
|
||
exists. The practical recipe that survives all controls: lens-locate the
|
||
band; anchor-merge with a state-dependent coefficient; label difficulty by
|
||
STaR; distill plans if you have them, loop if you don't; gate by predicted
|
||
difficulty; keep k=0 as the exact base model. What it buys: the
|
||
plan-dependent slice at zero tokens and zero decode cost. What it does not
|
||
buy: overall accuracy beyond what matched-compute sampling already delivers.
|
||
Both halves of that sentence are the contribution.
|