revision per review: seed means in headlines, net accounting vs untrained merge, bucket definition up front, GSM12B n.s. (p=0.86), law->regularity, norm spec, compute appendix, public repo URL

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Nils
2026-07-14 16:05:40 +02:00
co-authored by Claude Fable 5
parent 01fd272028
commit d1bafc2d4d
+139 -70
View File
@@ -1,8 +1,9 @@
# Latent Planning by Workspace Recurrence: an Interpretability-Placed Implant, and What It Actually Buys # Latent Planning by Workspace Recurrence: an Interpretability-Placed Implant, and What It Actually Buys
*Final-data draft, 2026-07-14. Base models: google/gemma-4-E2B-it and *Revision draft, 2026-07-14. Base models: google/gemma-4-E2B-it and
gemma-4-12B-it, both frozen. Hardware: DGX Spark + rented 2×/8×H100 nodes. gemma-4-12B-it, both frozen. Hardware: DGX Spark + rented 2×/8×H100 nodes.
Code, per-item logs, and pre-registrations: `~/jspace` (git). Statistics: Code, per-item logs, and pre-registrations:
https://git.draic.info/nils/jspace (public). Statistics:
`results-loop/STATS.md`.* `results-loop/STATS.md`.*
## Abstract ## Abstract
@@ -15,53 +16,56 @@ anchor-dominant merge adapter (0.03% of parameters) at the band entrance
turns the non-self-map band into a stable fixed-point iteration, trained with turns the non-self-map band into a stable fixed-point iteration, trained with
self-generated, verifier-filtered supervision. Looping the workspace over the self-generated, verifier-filtered supervision. Looping the workspace over the
prompt ("latent planning") raises pass@1 on plan-dependent MBPP problems from prompt ("latent planning") raises pass@1 on plan-dependent MBPP problems from
5.5% to 43.6% (seed mean 37.5±5.5), with zero visible tokens and zero 5.5% to 37.5±5.5 over five seeds (best seed 43.6%), with zero visible tokens
additional decode cost. The effect is real and highly reliable — pooled and zero additional decode cost. The effect is real and highly reliable —
across MBPP, HumanEval, and Rust/MultiPL-E, the plan-dependent bucket moves pooled across MBPP, HumanEval, and Rust/MultiPL-E, the plan-dependent bucket
from 4.2% to 35.6% (McNemar p≈1.5e-10) — and **placement is a law, not a moves from 4.2% to 35.6% (McNemar p≈1.5e-10) — and **placement is decisive,
convenience**: the gain appears only when the loop enters at the not a convenience**: the gain appears only when the loop enters at the
lens-identified boundary (L14), collapsing at L13 and below and structurally lens-identified boundary (L14), collapsing at L13 and below and structurally
nulling above. nulling above. Roughly a third of the gross effect is generic perturbation
(the untrained merge alone reaches ~18%; the bucket conditions on k=0
failure, so regression-to-mean contributes to any intervention); the
loop-specific net effect is ~+20 points over that floor.
But a complete attribution program deflates the mechanism's mystique. A complete attribution program then deflates the mechanism's mystique: the
(i) **The loop's content is amortizable**: distilling the model's own loop's content is *amortizable* (plan-distillation into the same adapter,
explicit plans into the same-size adapter — no recurrence at inference — recurrence-free, matches it; nothing stacks; looping distilled weights
matches or exceeds the loop on the same bucket (mean over 8 runs 45.7±4.6 vs degrades them), width rivals depth (trained pause registers reach 36.4%),
37.5±5.5, paired difference n.s.), and the two do not stack; running the and compute-matched token baselines win overall accuracy outright. What
loop on top of the distilled adapter *degrades* it. (ii) **Width rivals survives is precise: the implant owns exactly the plan-dependent slice at
depth**: 16 trained pause registers reach 36.4% on the same bucket. zero visible tokens and zero decode cost, transfers with the substrate
(iii) **Compute-matched token baselines are uncomfortable**: best-of-3
sampling beats every latent arm on overall accuracy (57.2% vs ≤55.2%), and a
50-token visible plan matches the loop on the hard bucket (40.0%). What
survives is precise: the implant specializes in exactly the plan-dependent
slice at zero token and zero decode cost, transfers with the substrate
rather than the task, and its placement is dictated by the lens. At 12B a rather than the task, and its placement is dictated by the lens. At 12B a
constant merge coefficient destroys the substrate; making the coefficient constant merge coefficient destroys the substrate; a state-dependent
state-dependent (a 3.8K-parameter gate) restores it on MBPP coefficient (3.8K parameters) restores MBPP but not Blocksworld or GSM8K —
(hard 11.4%→27.3% with overall preserved) but not on Blocksworld or GSM8K — the anchor coefficient is the stability dial that unifies this work with
the stability dial that unifies this work with McLeish et al. (2511.07384) McLeish et al. (2511.07384) and Lys et al. (2602.14759), and it is task-
and Lys et al. (2602.14759) is task- and scale-dependent. and scale-dependent. Details and exact numbers: §1 and §3.
## 1. What this paper claims ## 1. What this paper claims
1. **A placement law.** The retrofit works if and only if the recurrence *(One model family, two scales: we state findings as empirical regularities,
enters at the lens boundary. Entrances at L9L13 (same adapter, data, not laws.)*
curriculum) destroy overall accuracy (1434% vs 52%) while recovering at
most half the hard-bucket gain; entrance at L14 preserves overall and 1. **A placement regularity.** The retrofit works if and only if the
maximizes the gain (fig_placement). Entrances at L17/L24 are *structurally recurrence enters at the lens boundary. Entrances at L9L13 (same
null* in this architecture: KV-sharing makes layers ≥15 reuse keys/values adapter, data, curriculum) destroy overall accuracy (1434% vs 52%)
computed at ≤14, so k>0 is bit-identical to k=0 — a hazard for any while recovering at most half the hard-bucket gain; entrance at L14
retrofit method that skips the mechanistic check. Exit-layer choice is preserves overall and maximizes the gain (fig_placement). Entrances at
nearly free (taps 27/30/32/34 within seed noise: hard 3946%). This L17/L24 are *structurally null* in this architecture: KV-sharing makes
answers the open "where to loop" problem named by McLeish et al., and it layers ≥15 reuse keys/values computed at ≤14, so k>0 is bit-identical to
is causal, not correlational: the L9-entrance discriminator arm was k=0 — a hazard for any retrofit method that skips the mechanistic check.
trained identically and fails. Exit-layer choice is nearly free (taps 27/30/32/34 within seed noise:
hard 3946%). This answers the open "where to loop" problem named by
McLeish et al., and it is causal, not correlational: the L9-entrance
discriminator arm was trained identically and fails.
2. **A verified, statistically solid capability gain on a narrow slice.** 2. **A verified, statistically solid capability gain on a narrow slice.**
Plan-dependent items (the model solves them with an explicit written plan Plan-dependent items (the model solves them with an explicit written plan
but not directly): pooled across three benchmarks, 4.2%→35.6%, but not directly): seed-mean 37.5±5.5 on MBPP (best 43.6%); pooled across
p≈1.5e-10. Overall accuracy is statistically unchanged on MBPP three benchmarks, 4.2%→35.6%, p≈1.5e-10. Overall accuracy is
(p=0.34) and improved on HumanEval transfer (58.5%→66.5%, p=0.011). statistically unchanged on MBPP (p=0.34) and improved on HumanEval
transfer (58.5%→66.5%, p=0.011). Net of the untrained-merge floor
(~18%), the loop-specific effect is ~+20 points.
3. **A deflationary mechanism finding.** The trained loop converges to a 3. **A deflationary mechanism finding.** The trained loop converges to a
fixed point by k≈34 and behaves as *amortized plan content*, not fixed point by k≈34 and behaves as *amortized plan content*, not
@@ -73,7 +77,7 @@ and Lys et al. (2602.14759) is task- and scale-dependent.
adapter find plan-shaped content — content that can equally be put there adapter find plan-shaped content — content that can equally be put there
by distillation if plans are available. by distillation if plans are available.
4. **A width-vs-depth law.** Trained pause registers (width) capture most of 4. **A width-vs-depth pattern.** Trained pause registers (width) capture most of
the plan effect on code; recurrence (depth) is needed only where a state the plan effect on code; recurrence (depth) is needed only where a state
must *evolve* — on GSM8K generation-side carry beats registers, and on must *evolve* — on GSM8K generation-side carry beats registers, and on
Blocksworld (pure planning, no world knowledge) the loop lifts hard-split Blocksworld (pure planning, no world knowledge) the loop lifts hard-split
@@ -82,8 +86,10 @@ and Lys et al. (2602.14759) is task- and scale-dependent.
5. **Honest economics.** The implant's costs: ≈2.9× prompt-processing FLOPs 5. **Honest economics.** The implant's costs: ≈2.9× prompt-processing FLOPs
(parallel, prefill-shaped), zero decode overhead, bit-exact KV-cache (parallel, prefill-shaped), zero decode overhead, bit-exact KV-cache
write-in, k=0 recovers the base model exactly. Its competition at matched write-in, k=0 recovers the base model exactly. Its competition at
FLOPs: best-of-3 sampling wins overall accuracy outright (57.2%); a comparable compute (accounting in Appendix A — FLOPs, wall-clock, and
token budget do not rank the arms the same way): best-of-3 sampling wins
overall accuracy outright (57.2%); a
50-token visible plan ties the hard bucket. The value proposition is 50-token visible plan ties the hard bucket. The value proposition is
*only*: no visible tokens, no decode latency, and the hard-slice *only*: no visible tokens, no decode latency, and the hard-slice
specialization (distill's 46% > budget-CoT's 40% > best-of-3's 33%). specialization (distill's 46% > budget-CoT's 40% > best-of-3's 33%).
@@ -93,8 +99,8 @@ and Lys et al. (2602.14759) is task- and scale-dependent.
damage is present *before* adapter training (untrained-loop arm) and is damage is present *before* adapter training (untrained-loop arm) and is
not fixed by retuning α or LR. A per-position learned coefficient not fixed by retuning α or LR. A per-position learned coefficient
α=σ(w·[e;ŝ]+b) restores MBPP (overall 69.4%, hard 11.4%→27.3%) — but α=σ(w·[e;ŝ]+b) restores MBPP (overall 69.4%, hard 11.4%→27.3%) — but
fails to rescue Blocksworld-12B and yields only a marginal GSM8K-12B fails to rescue Blocksworld-12B and yields only a nominally positive,
overall gain (35.9%→36.7% at k=1), the project's only overall 12B win. not significant GSM8K-12B overall delta (35.9%→36.7% at k=1, p=0.86).
## 2. Method ## 2. Method
@@ -107,7 +113,9 @@ follow from what the readout tracks. On gemma-4-E2B: workspace ≈ L1430 of
(out-space ≠ in-space). With e = L13's output (fixed anchor) and s the (out-space ≠ in-space). With e = L13's output (fixed anchor) and s the
fed-back, norm-matched band output: fed-back, norm-matched band output:
L14-in = (1−α)·e + α·ŝ + MLP([e ; ŝ]), ŝ = s·|e|/|s| L14-in = (1−α)·e + α·ŝ + MLP([e ; ŝ]), ŝ = s·‖e‖₂/‖s‖₂
(per-position L2 norms over the hidden dimension, computed in fp32)
α=0.3 constant at 2B; at 12B, α=σ(w·[e;ŝ]+b) per position (zero-init so α=0.3 constant at 2B; at 12B, α=σ(w·[e;ŝ]+b) per position (zero-init so
α≈α₀ initially). MLP output zero-init: the untrained adapter is exactly the α≈α₀ initially). MLP output zero-init: the untrained adapter is exactly the
@@ -136,11 +144,23 @@ Cost at k=4: ≈2.9× prefill FLOPs, **zero** decode overhead.
Statistics throughout: Wilson 95% CIs; paired comparisons by exact McNemar; Statistics throughout: Wilson 95% CIs; paired comparisons by exact McNemar;
all headline arms evaluated on the full 500-item MBPP test split (hard all headline arms evaluated on the full 500-item MBPP test split (hard
bucket n=55), HumanEval n=164 (hard n=38), Rust/MultiPL-E n=154 (hard n=25), bucket n=55), HumanEval n=164 (hard n=38), Rust/MultiPL-E n=154 (hard n=25),
execution-verified. Label robustness: redefining "hard" as execution-verified.
labeled-hard ∧ k=0-fails-in-all-five-seeds (52/55 items) moves headline
numbers <2 points.
### 3.1 The placement law **Bucket definition, stated up front.** Hard labels come from labeling runs
of the frozen base model on the test items themselves (direct vs
plan-in-context, greedy). This is legitimate for *descriptive* slicing but
would be circular for selection — so no arm, hyperparameter, checkpoint, or
loop depth was ever chosen using bucket results (pre-registered;
`PROTOCOL_UNIFIED.md` items 12, 8). Because the bucket conditions on k=0
failure, regression-to-mean inflates *any* intervention's bucket score: the
untrained merge already reaches ~18%, and we therefore report the
loop-specific effect **net of that floor** wherever attribution is claimed.
Robustness: redefining "hard" as labeled-hard ∧ k=0-fails-in-all-five-seeds
(52/55 items) moves headline numbers <2 points; both definitions share the
base model, which an independent difficulty proxy would not — we flag this
as an open external check.
### 3.1 The placement regularity
![Placement cliff](results-loop/fig_placement.png) ![Placement cliff](results-loop/fig_placement.png)
@@ -150,8 +170,8 @@ hard 43.6%, overall 53.6%. L13: hard 17.9%, overall 34.4%. L9L12: overall
verified bit-identical) — the 12B model has no shared-KV layers, making it verified bit-identical) — the 12B model has no shared-KV layers, making it
the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34): the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34):
hard 39.346.4%, within seed spread. The lens boundary is necessary; the hard 39.346.4%, within seed spread. The lens boundary is necessary; the
exit is a free parameter. (The L23-exit arm died in training and was not exit is a free parameter. (The L23-exit arm died in training; a rerun is in
rerun; the four completed exits bracket it.) progress — the four completed exits bracket it. [L23 PENDING])
### 3.2 The attribution ladder ### 3.2 The attribution ladder
@@ -178,6 +198,14 @@ vs stack) is **not significant** at n=55; loop vs base *overall* is not
significant on MBPP (p=0.34). The ladder's shape is reliable; its fine significant on MBPP (p=0.34). The ladder's shape is reliable; its fine
ordering is not. ordering is not.
**Net accounting.** The attribution-critical comparison is trained-loop vs
*untrained merge*, not vs base: gross 5.5→37.5 (seed mean), of which the
untrained perturbation floor is ~18 points — the loop-specific net effect
is ~+20 points. The untrained-loop and trained-FF control rows above are
from the 250-item era (hard n=28); full-bucket (n=55) reruns of both
controls, enabling the paired loop-vs-untrained test, are running and will
replace these rows. [CONTROLS-N55 PENDING]
### 3.3 The decisive tests: nothing stacks ### 3.3 The decisive tests: nothing stacks
If the loop performed genuine iterative computation, plan-distilled content If the loop performed genuine iterative computation, plan-distilled content
@@ -199,17 +227,21 @@ distinctive value is that it finds this content *without* plan supervision
### 3.4 Compute-matched honesty ### 3.4 Compute-matched honesty
At approximately matched FLOPs, token-space baselines are strong: best-of-3 At approximately matched FLOPs (Appendix A gives the accounting, separated
sampling wins overall accuracy against every latent arm (57.2%, into FLOPs, wall-clock, and token budget), token-space baselines are strong:
best-of-3 sampling wins overall accuracy against every latent arm (57.2%,
CI [52.8, 61.5], vs loop 53.6 [49.2, 57.9] — point estimate higher, CIs CI [52.8, 61.5], vs loop 53.6 [49.2, 57.9] — point estimate higher, CIs
overlap) by preserving easy items perfectly while sampling rescues some hard overlap) by preserving easy items perfectly while sampling rescues some hard
ones. A 50-token visible plan ties the loop's hard bucket. The latent ones. A 50-token visible plan ties the loop's hard bucket. Both baselines
implant's surviving advantages are qualitative: zero visible tokens (silent), are being rerun with per-item logs to enable paired tests against the latent
zero decode overhead (prefill-parallel; sampling and CoT pay serially at arms; until those land, the overall-accuracy comparison rests on overlapping
bandwidth-bound decode), and the hard-slice crown under distillation (46% vs CIs and is stated as point-estimate-level. [BASELINES-PI PENDING]
40% budget-CoT vs 33% best-of-3). For deployment this means: the implant is The latent implant's surviving advantages are qualitative: zero visible
a *latency/token-budget* technology with a side specialization in tokens (silent), zero decode overhead (prefill-parallel; sampling and CoT
plan-dependent items — not an accuracy technology. pay serially at bandwidth-bound decode), and the hard-slice crown under
distillation (46% vs 40% budget-CoT vs 33% best-of-3). For deployment this
means: the implant is a *latency/token-budget* technology with a side
specialization in plan-dependent items — not an accuracy technology.
### 3.5 Width vs depth, and the task boundary ### 3.5 Width vs depth, and the task boundary
@@ -219,7 +251,7 @@ prompt-side result entirely (no variant beats the weights control
prompt-side), but generation-side *carry* — recurrence across token steps — prompt-side), but generation-side *carry* — recurrence across token steps —
doubles the pause control on hard items: arithmetic's serial state evolves doubles the pause control on hard items: arithmetic's serial state evolves
during the answer. Blocksworld at 2B is the purest case: base 0% on hard during the answer. Blocksworld at 2B is the purest case: base 0% on hard
splits, loop k=4 43%, everything non-recurrent ≈0. The law: **plans are splits, loop k=4 43%, everything non-recurrent ≈0. The pattern: **plans are
wide; execution is deep.** Retrofit recurrence pays off precisely where a wide; execution is deep.** Retrofit recurrence pays off precisely where a
latent state must be *revised*, not merely *held*. latent state must be *revised*, not merely *held*.
@@ -233,9 +265,11 @@ the damage precedes adapter training; α=0.15 and LR retuning do not fix it
(47.6/52.6% overall). The state-dependent coefficient does, on MBPP: (47.6/52.6% overall). The state-dependent coefficient does, on MBPP:
overall 69.4% (base 72.4%), hard 11.4%→27.3%. It does **not** rescue overall 69.4% (base 72.4%), hard 11.4%→27.3%. It does **not** rescue
Blocksworld-12B (easy items destroyed at k=4; constant-α had reached hard Blocksworld-12B (easy items destroyed at k=4; constant-α had reached hard
40% but also destroyed easy) and yields only +0.8 points overall on 40% but also destroyed easy) and yields a **nominally positive, not
GSM8K-12B (35.9→36.7 at k=1, hard 1.6→10.6) — the sole overall-accuracy win significant** overall delta on GSM8K-12B (35.9→36.7 at k=1; paired McNemar
of the program, and a marginal one. Conclusion: the anchor coefficient is on 32 discordant items, p=0.86; hard 1.6→10.6) — no arm anywhere in the
program produced a statistically significant overall gain at 12B.
Conclusion: the anchor coefficient is
the load-bearing stability control, its correct *form* (not just value) the load-bearing stability control, its correct *form* (not just value)
changes with scale, and per-task tuning remains unavoidable. changes with scale, and per-task tuning remains unavoidable.
@@ -314,7 +348,7 @@ proved a genuine rival, not a strawman (§3.2, §3.5).
What remains distinct here: interpretability-derived placement with causal What remains distinct here: interpretability-derived placement with causal
validation; a fully frozen base with bit-exact k=0 and zero-decode-cost KV validation; a fully frozen base with bit-exact k=0 and zero-decode-cost KV
write-in; the complete attribution ladder including compute-matched write-in; the complete attribution ladder including compute-matched
token-space baselines and stacking tests; the width/depth task law; and the token-space baselines and stacking tests; the width/depth task pattern; and the
amortizability finding itself. amortizability finding itself.
## 5. Limitations ## 5. Limitations
@@ -324,10 +358,9 @@ are small (n=55/38/25); within-ladder orderings are not individually
significant, and only the pooled hard effect and the HumanEval overall gain significant, and only the pooled hard effect and the HumanEval overall gain
survive multiple-comparison scrutiny. Bucket membership derives from greedy survive multiple-comparison scrutiny. Bucket membership derives from greedy
labeling runs (consensus-k0 robustness check moves numbers <2 points, but labeling runs (consensus-k0 robustness check moves numbers <2 points, but
both checks share the base model). Best-of-3/budget-CoT lack per-item logs both checks share the base model; an independent difficulty proxy is an
(no paired tests against them). The L23 exit arm and a third architecture open external check). A third architecture family was not run;
family were not run; LiveCodeBench (contamination-safe) was not run; rung-2 LiveCodeBench (contamination-safe) was not run; rung-2 was not run at 12B. The easy-item perturbation tax persists wherever the
was not run at 12B. The easy-item perturbation tax persists wherever the
gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both
arms share contamination, and memorized items land in the easy bucket, but arms share contamination, and memorized items land in the easy bucket, but
bucket composition is contamination-sensitive. The capability panel bucket composition is contamination-sensitive. The capability panel
@@ -351,3 +384,39 @@ difficulty; keep k=0 as the exact base model. What it buys: the
plan-dependent slice at zero tokens and zero decode cost. What it does not plan-dependent slice at zero tokens and zero decode cost. What it does not
buy: overall accuracy beyond what matched-compute sampling already delivers. buy: overall accuracy beyond what matched-compute sampling already delivers.
Both halves of that sentence are the contribution. Both halves of that sentence are the contribution.
## Appendix A: compute accounting (FLOPs / wall-clock / tokens, separated)
Let P = prompt tokens, G = generated tokens, c = FLOPs per token per full
forward pass. The band is 17 of 35 decoder layers at E2B (fraction
f≈0.486) and 10 of 48 at 12B (f≈0.208).
**Latent loop, k=4, prompt-only.** Prefill: 1 base pass + 4 band passes
over prompt positions = (1+4f)·cP ≈ **2.94·cP** at E2B (1.83× at 12B —
the overhead *shrinks* with scale because lens bands grow sublinearly).
Decode: exactly cG (looped states written to the KV cache once;
bit-exactness verified). Wall-clock: prefill is compute-bound and
position-parallel, but the k iterations are serial — prefill latency
≈2.9×, typically a small fraction of end-to-end latency for G≫0.
Visible tokens: +0.
**Best-of-3 sampling.** FLOPs: with shared prompt prefill (favorable
accounting), cP + 3·cG ≈ cP + 3cG; without sharing 3c(P+G). For MBPP
(P≈150300, G≈150220), the *extra* FLOPs vs direct (≈2cG) are of the same
order as the loop's extra (≈1.94cP) — hence "≈matched". Wall-clock: 3×G
serial bandwidth-bound decode steps (or 3 parallel decode streams at 3×
memory); strictly worse latency than the loop unless parallelized.
Visible tokens: ≈3× (two discarded candidates). Requires a verifier or
selector to pick among samples for the overall win we report (we use
any-pass, an upper bound — see §3.4 caveat).
**Budget-CoT (50-token plan).** FLOPs: ≈c(P+G+50) plus the plan tokens'
KV in context for the remainder — the *cheapest* arm in FLOPs. Wall-clock:
+50 serial decode steps before answer tokens start (worst first-token
latency). Visible tokens: +50.
Summary: no single scalar makes these three arms "equal"; the loop
dominates on tokens and decode latency, budget-CoT on FLOPs, best-of-3 on
overall accuracy. §3.4's "≈matched FLOPs" refers to the extra-FLOPs
order-of-magnitude equivalence above, not exact equality; the honest
statement is the three-way trade-off, and we report all three axes.