426 lines
24 KiB
Markdown
426 lines
24 KiB
Markdown
# Latent Planning by Workspace Recurrence: an Interpretability-Placed Implant, and What It Actually Buys
|
||
|
||
*Revision draft, 2026-07-14. Base models: google/gemma-4-E2B-it and
|
||
gemma-4-12B-it, both frozen. Hardware: DGX Spark + rented 2×/8×H100 nodes.
|
||
Code, per-item logs, and pre-registrations:
|
||
https://git.draic.info/nils/jspace (public). Statistics:
|
||
`results-loop/STATS.md`.*
|
||
|
||
## Abstract
|
||
|
||
Interpretability work with an averaged-Jacobian lens ("J-lens") partitions a
|
||
pretrained language model's depth into regimes, including a mid-depth
|
||
*workspace* band that holds verbalizable, unspoken intermediate content. We
|
||
retrofit recurrence onto this band in a **frozen** model: a 1.6M-parameter
|
||
anchor-dominant merge adapter (0.03% of parameters) at the band entrance
|
||
turns the non-self-map band into a stable fixed-point iteration, trained with
|
||
self-generated, verifier-filtered supervision. Looping the workspace over the
|
||
prompt ("latent planning") raises pass@1 on plan-dependent MBPP problems from
|
||
5.5% to 37.5±5.5 over five seeds (best seed 43.6%), with zero visible tokens
|
||
and zero additional decode cost. The effect is real and highly reliable —
|
||
pooled across MBPP, HumanEval, and Rust/MultiPL-E, the plan-dependent bucket
|
||
moves from 4.2% to 35.6% (McNemar p≈1.5e-10) — and **placement is decisive,
|
||
not a convenience**: the gain appears only when the loop enters at the
|
||
lens-identified boundary (L14), collapsing at L13 and below and structurally
|
||
nulling above. Roughly a third of the gross effect is generic perturbation
|
||
(the untrained merge alone reaches ~18%; the bucket conditions on k=0
|
||
failure, so regression-to-mean contributes to any intervention); the
|
||
loop-specific net effect is ~+20 points over that floor.
|
||
|
||
A complete attribution program then deflates the mechanism's mystique: the
|
||
loop's content is *amortizable* (plan-distillation into the same adapter,
|
||
recurrence-free, matches it; nothing stacks; looping distilled weights
|
||
degrades them), width rivals depth (trained pause registers reach 36.4%),
|
||
and compute-matched token baselines win overall accuracy outright. What
|
||
survives is precise: the implant owns exactly the plan-dependent slice at
|
||
zero visible tokens and zero decode cost, transfers with the substrate
|
||
rather than the task, and its placement is dictated by the lens. At 12B a
|
||
constant merge coefficient destroys the substrate; a state-dependent
|
||
coefficient (3.8K parameters) restores MBPP but not Blocksworld or GSM8K —
|
||
the anchor coefficient is the stability dial that unifies this work with
|
||
McLeish et al. (2511.07384) and Lys et al. (2602.14759), and it is task-
|
||
and scale-dependent. Details and exact numbers: §1 and §3.
|
||
|
||
## 1. What this paper claims
|
||
|
||
*(One model family, two scales: we state findings as empirical regularities,
|
||
not laws.)*
|
||
|
||
1. **A placement regularity.** The retrofit works if and only if the
|
||
recurrence enters at the lens boundary. Entrances at L9–L13 (same
|
||
adapter, data, curriculum) destroy overall accuracy (14–34% vs 52%)
|
||
while recovering at most half the hard-bucket gain; entrance at L14
|
||
preserves overall and maximizes the gain (fig_placement). Entrances at
|
||
L17/L24 are *structurally null* in this architecture: KV-sharing makes
|
||
layers ≥15 reuse keys/values computed at ≤14, so k>0 is bit-identical to
|
||
k=0 — a hazard for any retrofit method that skips the mechanistic check.
|
||
Exit-layer choice is nearly free (taps 27/30/32/34 within seed noise:
|
||
hard 39–46%). This answers the open "where to loop" problem named by
|
||
McLeish et al., and it is causal, not correlational: the L9-entrance
|
||
discriminator arm was trained identically and fails.
|
||
|
||
2. **A verified, statistically solid capability gain on a narrow slice.**
|
||
Plan-dependent items (the model solves them with an explicit written plan
|
||
but not directly): seed-mean 37.5±5.5 on MBPP (best 43.6%); pooled across
|
||
three benchmarks, 4.2%→35.6%, p≈1.5e-10. Overall accuracy is
|
||
statistically unchanged on MBPP (p=0.34) and improved on HumanEval
|
||
transfer (58.5%→66.5%, p=0.011). Net of the untrained-merge floor
|
||
(~18%), the loop-specific effect is ~+20 points.
|
||
|
||
3. **A deflationary mechanism finding.** The trained loop converges to a
|
||
fixed point by k≈3–4 and behaves as *amortized plan content*, not
|
||
iterative computation: plan-distillation into the identical architecture
|
||
without recurrence matches it; stacking buys nothing (loop-training a
|
||
distill-warmed adapter: 34.5%, below distill alone; running the distilled
|
||
adapter in loop mode: drops to 20.0%); deeper k at inference is flat
|
||
(k=8: 40.0%). The recurrence is a *training-time scaffold* that lets the
|
||
adapter find plan-shaped content — content that can equally be put there
|
||
by distillation if plans are available.
|
||
|
||
4. **A width-vs-depth pattern.** Trained pause registers (width) capture most of
|
||
the plan effect on code; recurrence (depth) is needed only where a state
|
||
must *evolve* — on GSM8K generation-side carry beats registers, and on
|
||
Blocksworld (pure planning, no world knowledge) the loop lifts hard-split
|
||
plans 0%→43% at 2B where everything else fails. Plans are wide; execution
|
||
is deep.
|
||
|
||
5. **Honest economics.** The implant's costs: ≈2.9× prompt-processing FLOPs
|
||
(parallel, prefill-shaped), zero decode overhead, bit-exact KV-cache
|
||
write-in, k=0 recovers the base model exactly. Its competition at
|
||
comparable compute (accounting in Appendix A — FLOPs, wall-clock, and
|
||
token budget do not rank the arms the same way): best-of-3 sampling wins
|
||
overall accuracy outright (57.2%); a
|
||
50-token visible plan ties the hard bucket. The value proposition is
|
||
*only*: no visible tokens, no decode latency, and the hard-slice
|
||
specialization (distill's 46% > budget-CoT's 40% > best-of-3's 33%).
|
||
|
||
6. **Scale transfers only with a state-dependent stability dial.** At 12B the
|
||
2B-tuned constant α=0.3 collapses overall accuracy (72.6%→43.0%); the
|
||
damage is present *before* adapter training (untrained-loop arm) and is
|
||
not fixed by retuning α or LR. A per-position learned coefficient
|
||
α=σ(w·[e;ŝ]+b) restores MBPP (overall 69.4%, hard 11.4%→27.3%) — but
|
||
fails to rescue Blocksworld-12B and yields only a nominally positive,
|
||
not significant GSM8K-12B overall delta (35.9%→36.7% at k=1, p=0.86).
|
||
|
||
## 2. Method
|
||
|
||
**Locating the band.** The lens reads residual state h at layer ℓ through the
|
||
averaged Jacobian J̄_ℓ = E[∂h_final/∂h_ℓ] and the unembedding; depth regimes
|
||
follow from what the readout tracks. On gemma-4-E2B: workspace ≈ L14–30 of
|
||
35; on 12B: L36–45 of 48.
|
||
|
||
**Making the band a self-map.** Feeding L30's output back to L14 collapses
|
||
(out-space ≠ in-space). With e = L13's output (fixed anchor) and s the
|
||
fed-back, norm-matched band output:
|
||
|
||
L14-in = (1−α)·e + α·ŝ + MLP([e ; ŝ]), ŝ = s·‖e‖₂/‖s‖₂
|
||
|
||
(per-position L2 norms over the hidden dimension, computed in fp32)
|
||
|
||
α=0.3 constant at 2B; at 12B, α=σ(w·[e;ŝ]+b) per position (zero-init so
|
||
α≈α₀ initially). MLP output zero-init: the untrained adapter is exactly the
|
||
hand merge — stable, answer-preserving, content-holding.
|
||
|
||
**Training.** STaR-style self-labeling: items the frozen model solves only
|
||
with an explicit plan/CoT are "hard", direct solves "easy", neither "drop".
|
||
Cross-entropy on answer/code tokens of the direct prompt, model's own
|
||
verified outputs as targets. Difficulty→depth curriculum (easy k=1, mixed
|
||
k=2, hard k=2–4). The loop applies to the **prompt span only**; generated
|
||
tokens run the plain path but attend to looped prompt states. Variants
|
||
trained the same way: **pause-N** (N trained register tokens appended to the
|
||
prompt, no recurrence), **plan-distill** (KL from the model's own
|
||
plan-in-context distribution into the FF adapter), **rung-2** (warm-started
|
||
adapter + entrance-faded LoRA rank 8 on the band's first layers, loop-only
|
||
via a global toggle), and **stack** arms (distill-warm + loop training;
|
||
distilled adapter evaluated in loop mode).
|
||
|
||
**Inference.** Looped prompt states are causally independent of generated
|
||
tokens: computed once at prefill, written into the KV cache by a hooked
|
||
forward pass, generation native. Verified bit-identical to the slow path.
|
||
Cost at k=4: ≈2.9× prefill FLOPs, **zero** decode overhead.
|
||
|
||
## 3. Results
|
||
|
||
Statistics throughout: Wilson 95% CIs; paired comparisons by exact McNemar;
|
||
all headline arms evaluated on the full 500-item MBPP test split (hard
|
||
bucket n=55), HumanEval n=164 (hard n=38), Rust/MultiPL-E n=154 (hard n=25),
|
||
execution-verified.
|
||
|
||
**Bucket definition, stated up front.** Hard labels come from labeling runs
|
||
of the frozen base model on the test items themselves (direct vs
|
||
plan-in-context, greedy). This is legitimate for *descriptive* slicing but
|
||
would be circular for selection — so no arm, hyperparameter, checkpoint, or
|
||
loop depth was ever chosen using bucket results (pre-registered;
|
||
`PROTOCOL_UNIFIED.md` items 1–2, 8). Because the bucket conditions on k=0
|
||
failure, regression-to-mean inflates *any* intervention's bucket score: the
|
||
untrained merge already reaches ~18%, and we therefore report the
|
||
loop-specific effect **net of that floor** wherever attribution is claimed.
|
||
Robustness: redefining "hard" as labeled-hard ∧ k=0-fails-in-all-five-seeds
|
||
(52/55 items) moves headline numbers <2 points; both definitions share the
|
||
base model, which an independent difficulty proxy would not — we flag this
|
||
as an open external check.
|
||
|
||
### 3.1 The placement regularity
|
||
|
||

|
||
|
||
Entrance-layer sweep with everything else fixed. L14 (lens boundary):
|
||
hard 43.6%, overall 53.6%. L13: hard 17.9%, overall 34.4%. L9–L12: overall
|
||
14.0–30.8% (substrate destroyed). L17/L24 entrances: k>0 ≡ k=0 (KV sharing;
|
||
verified bit-identical) — the 12B model has no shared-KV layers, making it
|
||
the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34):
|
||
hard 39.3–46.4%, within seed spread. The lens boundary is necessary; the
|
||
exit is a free parameter. (The L23-exit arm died in training; a rerun is in
|
||
progress — the four completed exits bracket it. [L23 PENDING])
|
||
|
||
### 3.2 The attribution ladder
|
||
|
||

|
||
|
||
MBPP hard bucket (plan-dependent, n=55 unless noted):
|
||
|
||
| arm | hard pass@1 | overall |
|
||
|---|---|---|
|
||
| base (k=0, bit-exact) | 5.5% | 51.8% |
|
||
| untrained loop (α-merge only, n=28) | 17.9% | ~52% |
|
||
| trained FF, no recurrence (n=28) | 17.9% | ~52% |
|
||
| pause-16 registers (width) | 36.4% | 55.2% |
|
||
| **trained loop k=4** (seed mean, 5 seeds) | **37.5±5.5** (best 43.6) | 53.6% |
|
||
| rung-2: + entrance-faded band LoRA (n=28) | 42.9/46.4 (2 seeds) | 51.2/52.4 |
|
||
| **plan-distilled FF** (mean, 8 runs) | **45.7±4.6** (best 49.1) | 55.5% |
|
||
| budget-CoT (50 visible tokens) | 40.0% | 53.8% |
|
||
| best-of-3 sampling (≈matched FLOPs) | 32.7% | **57.2%** |
|
||
| explicit plan in context (ceiling) | 94.5% | 59.0% |
|
||
|
||
Significance structure (McNemar, `STATS.md`): loop vs base on hard,
|
||
p=5.7e-6; every latent-arm-vs-latent-arm difference (loop vs distill, distill
|
||
vs stack) is **not significant** at n=55; loop vs base *overall* is not
|
||
significant on MBPP (p=0.34). The ladder's shape is reliable; its fine
|
||
ordering is not.
|
||
|
||
**Net accounting.** The attribution-critical comparison is trained-loop vs
|
||
*untrained merge*, not vs base: gross 5.5→37.5 (seed mean), of which the
|
||
untrained perturbation floor is ~18 points — the loop-specific net effect
|
||
is ~+20 points. The untrained-loop and trained-FF control rows above are
|
||
from the 250-item era (hard n=28); full-bucket (n=55) reruns of both
|
||
controls, enabling the paired loop-vs-untrained test, are running and will
|
||
replace these rows. [CONTROLS-N55 PENDING]
|
||
|
||
### 3.3 The decisive tests: nothing stacks
|
||
|
||
If the loop performed genuine iterative computation, plan-distilled content
|
||
plus recurrence should compound. It does not:
|
||
|
||
- **Distill-warm + loop training**: hard 34.5% — below distill alone.
|
||
- **Distilled adapter run in loop mode**: hard 20.0%, overall 45.8% —
|
||
looping *degrades* the distilled weights.
|
||
- **Pause-16 + distill**: hard 30.9% — no width stacking either.
|
||
- **Inference depth beyond convergence**: k=8 hard 40.0% ≈ k=4 (fixed point,
|
||
cos(sₖ,sₖ₋₁)=1.000 by k≈3–4).
|
||
|
||
Reading: the recurrence is a **training-time scaffold**. The curriculum
|
||
forces hard-item loss to be reducible only through the loop, and what the
|
||
adapter learns to inject is plan-shaped content — the same content
|
||
distillation installs directly when explicit plans are available. The loop's
|
||
distinctive value is that it finds this content *without* plan supervision
|
||
(STaR labels only say which items needed plans, not what the plans were).
|
||
|
||
### 3.4 Compute-matched honesty
|
||
|
||
At approximately matched FLOPs (Appendix A gives the accounting, separated
|
||
into FLOPs, wall-clock, and token budget), token-space baselines are strong:
|
||
best-of-3 sampling wins overall accuracy against every latent arm (57.2%,
|
||
CI [52.8, 61.5], vs loop 53.6 [49.2, 57.9] — point estimate higher, CIs
|
||
overlap) by preserving easy items perfectly while sampling rescues some hard
|
||
ones. A 50-token visible plan ties the loop's hard bucket. Both baselines
|
||
are being rerun with per-item logs to enable paired tests against the latent
|
||
arms; until those land, the overall-accuracy comparison rests on overlapping
|
||
CIs and is stated as point-estimate-level. [BASELINES-PI PENDING]
|
||
The latent implant's surviving advantages are qualitative: zero visible
|
||
tokens (silent), zero decode overhead (prefill-parallel; sampling and CoT
|
||
pay serially at bandwidth-bound decode), and the hard-slice crown under
|
||
distillation (46% vs 40% budget-CoT vs 33% best-of-3). For deployment this
|
||
means: the implant is a *latency/token-budget* technology with a side
|
||
specialization in plan-dependent items — not an accuracy technology.
|
||
|
||
### 3.5 Width vs depth, and the task boundary
|
||
|
||
Pause registers (width) reach 36.4% (16 registers; 8: 30.9%, 32: 34.5% — flat
|
||
in N) on MBPP hard: static plan content fits in registers. GSM8K inverts the
|
||
prompt-side result entirely (no variant beats the weights control
|
||
prompt-side), but generation-side *carry* — recurrence across token steps —
|
||
doubles the pause control on hard items: arithmetic's serial state evolves
|
||
during the answer. Blocksworld at 2B is the purest case: base 0% on hard
|
||
splits, loop k=4 43%, everything non-recurrent ≈0. The pattern: **plans are
|
||
wide; execution is deep.** Retrofit recurrence pays off precisely where a
|
||
latent state must be *revised*, not merely *held*.
|
||
|
||
### 3.6 Scale: the stability dial
|
||
|
||

|
||
|
||
At 12B (no shared KV — unconfounded), constant α=0.3: overall collapses
|
||
72.6%→43.0% at k=4 while hard limps to 11.4%. The untrained-loop arm shows
|
||
the damage precedes adapter training; α=0.15 and LR retuning do not fix it
|
||
(47.6/52.6% overall). The state-dependent coefficient does, on MBPP:
|
||
overall 69.4% (base 72.4%), hard 11.4%→27.3%. It does **not** rescue
|
||
Blocksworld-12B (easy items destroyed at k=4; constant-α had reached hard
|
||
40% but also destroyed easy) and yields a **nominally positive, not
|
||
significant** overall delta on GSM8K-12B (35.9→36.7 at k=1; paired McNemar
|
||
on 32 discordant items, p=0.86; hard 1.6→10.6) — no arm anywhere in the
|
||
program produced a statistically significant overall gain at 12B.
|
||
Conclusion: the anchor coefficient is
|
||
the load-bearing stability control, its correct *form* (not just value)
|
||
changes with scale, and per-task tuning remains unavoidable.
|
||
|
||
### 3.7 Transfer: substrate, not task
|
||
|
||

|
||
|
||
MBPP-trained implants applied unchanged: **HumanEval** overall 58.5%→66.5%
|
||
(loop k=4, p=0.011 vs base; hard 0→31.6%). The decisive control: the
|
||
*untrained* merge already reaches 64.6%, and trained-vs-untrained is **not
|
||
significant** (paired McNemar at k=2, 9 vs 7 discordant, p=0.80). What
|
||
transfers significantly is the *merge perturbation itself*, not the
|
||
MBPP-trained content — the cleanest evidence that off-distribution value is
|
||
substrate-shaped rather than task-memorized. (The transferred pause adapter
|
||
reaches 66.5%, hard 38.9%, consistent with the same reading.) **Rust/MultiPL-E** (Python-trained, different
|
||
language, compile-run-verified): hard 8.0%→24.0% (p=0.125 at n=25 —
|
||
directionally consistent, underpowered). **Blocksworld** MBPP-transfer:
|
||
hard 0→14.3% (task-trained: 43%). Content transfers where the substrate's
|
||
plan-representation overlaps; task-specific training still dominates.
|
||
|
||
### 3.8 Mechanism, verification, deployment
|
||
|
||
The trained loop takes a large first step (cos(s₁,s₀)=0.926 vs 0.977
|
||
untrained) and converges bit-exactly by k≈3–4; accuracy and lens-sharpening
|
||
plateau there. P(latent concept) under the J-lens at the band exit rises
|
||
0.015→0.13 across iterations (~8× the untrained hold) — the lens that placed
|
||
the implant also renders its silent content inspectable. The STaR labels
|
||
train a free difficulty gate (route predicted-hard to k=4, else k=0);
|
||
gate quality (19% precision at 64% recall) is the current ceiling on
|
||
removing the easy-item perturbation tax. k=0 is the exact base model by
|
||
construction — the implant is removable at token granularity.
|
||
|
||
**General-capability panel** (ARC-Challenge, WinoGrande, HellaSwag, MMLU;
|
||
length-normalized MC scoring with the loop applied to the context span) is
|
||
running on the Spark; results will quantify what k>0 does to off-task
|
||
abilities. [PENDING — fill on completion.]
|
||
|
||
### 3.9 Negative results with content
|
||
|
||
Mixed-task (code+math) training regressed both tasks at equal validation CE
|
||
— CE parity does not predict generation parity, and validation-CE checkpoint
|
||
selection fails likewise (fixed-step pre-commitment used instead; no
|
||
checkpoint was selected on test or generation results). GSM8K distillation
|
||
collapsed to empty outputs twice (E2B first attempt, 12B) on 3-token targets
|
||
under KL-dominant loss; a CE-dominant retry at E2B trained but reached only
|
||
hard 4.7%. Plan-distillation on GSM8K underperforms its MBPP twin even when
|
||
training succeeds: consistent with §3.5, there is little static plan content
|
||
for math to amortize.
|
||
|
||
## 4. Related work
|
||
|
||
**McLeish et al. (arXiv:2511.07384)** retrofit depth-recurrence via layer
|
||
surgery + ~50B-token continued pretraining of all parameters; they name
|
||
layer choice as an open problem — §3.1 is a causal answer. Their surgery
|
||
needs a healing phase; our k=0 is exactly the base model. **Lys et al.
|
||
(arXiv:2602.14759)** loop frozen models training-free; their finding that
|
||
naive looping degrades while interpolation with the un-looped state rescues
|
||
it is independent convergent evidence for anchor-dominance, and their
|
||
setting is the untrained cell of our ladder (17.9%).
|
||
|
||
**One mechanism, three regimes.** All three works mix the fed-back state
|
||
with an anchor from the un-looped computation. Lys et al.'s moving average
|
||
η·h⁽⁰⁾+(1−η)·h⁽ᵗ⁾ is an untrained anchor coefficient; our
|
||
(1−α)e + α·ŝ + MLP is its trained analogue; McLeish et al.'s input injection
|
||
is the fully-learned limit. The 12B episode closes the loop on this
|
||
unification: the coefficient is the stability dial, naive looping is its
|
||
α→1 collapse limit, and our scale failure + state-dependent fix show the
|
||
dial must itself become a function of the state as models grow. Our stacking
|
||
results add a caution for the whole family: if retrofitted recurrence
|
||
content is amortizable (§3.3), some of the family's gains may be
|
||
reproducible by distillation without inference-time recurrence — a control
|
||
neither bracket paper runs.
|
||
|
||
Earlier lineage: Universal Transformers; DEQ; Huginn (2502.05171);
|
||
Mixture-of-Recursions (2507.10524); Relaxed Recursive Transformers
|
||
(2410.20672); Coconut; pause tokens (Goyal et al.) — whose trained variant
|
||
proved a genuine rival, not a strawman (§3.2, §3.5).
|
||
|
||
What remains distinct here: interpretability-derived placement with causal
|
||
validation; a fully frozen base with bit-exact k=0 and zero-decode-cost KV
|
||
write-in; the complete attribution ladder including compute-matched
|
||
token-space baselines and stacking tests; the width/depth task pattern; and the
|
||
amortizability finding itself.
|
||
|
||
## 5. Limitations
|
||
|
||
One model family (gemma-4), two scales, three task families. Hard buckets
|
||
are small (n=55/38/25); within-ladder orderings are not individually
|
||
significant, and only the pooled hard effect and the HumanEval overall gain
|
||
survive multiple-comparison scrutiny. Bucket membership derives from greedy
|
||
labeling runs (consensus-k0 robustness check moves numbers <2 points, but
|
||
both checks share the base model; an independent difficulty proxy is an
|
||
open external check). A third architecture family was not run;
|
||
LiveCodeBench (contamination-safe) was not run; rung-2 was not run at 12B. The easy-item perturbation tax persists wherever the
|
||
gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both
|
||
arms share contamination, and memorized items land in the easy bucket, but
|
||
bucket composition is contamination-sensitive. The capability panel
|
||
(§3.8) is pending; until it lands, off-task effects of k>0 are unmeasured.
|
||
The Blocksworld-12B and GSM8K-12B failures mean the adaptive-α fix is
|
||
demonstrated on one task at one scale, not established as a general recipe.
|
||
|
||
## 6. Conclusion
|
||
|
||
The experiment this program set out to run — *can an interpretability lens
|
||
tell you where to install recurrence in a frozen model, and does it work?* —
|
||
has a clean answer: yes, and the placement is causally load-bearing. The
|
||
more interesting answer is what the recurrence turned out to be: not a
|
||
reasoning engine, but a remarkably cheap way to make a frozen model amortize
|
||
its own planning into 0.03% of extra parameters, with a training-time loop
|
||
as scaffold and an inference-time loop that is optional once the content
|
||
exists. The practical recipe that survives all controls: lens-locate the
|
||
band; anchor-merge with a state-dependent coefficient; label difficulty by
|
||
STaR; distill plans if you have them, loop if you don't; gate by predicted
|
||
difficulty; keep k=0 as the exact base model. What it buys: the
|
||
plan-dependent slice at zero tokens and zero decode cost. What it does not
|
||
buy: overall accuracy beyond what matched-compute sampling already delivers.
|
||
Both halves of that sentence are the contribution.
|
||
|
||
## Appendix A: compute accounting (FLOPs / wall-clock / tokens, separated)
|
||
|
||
Let P = prompt tokens, G = generated tokens, c = FLOPs per token per full
|
||
forward pass. The band is 17 of 35 decoder layers at E2B (fraction
|
||
f≈0.486) and 10 of 48 at 12B (f≈0.208).
|
||
|
||
**Latent loop, k=4, prompt-only.** Prefill: 1 base pass + 4 band passes
|
||
over prompt positions = (1+4f)·cP ≈ **2.94·cP** at E2B (1.83× at 12B —
|
||
the overhead *shrinks* with scale because lens bands grow sublinearly).
|
||
Decode: exactly cG (looped states written to the KV cache once;
|
||
bit-exactness verified). Wall-clock: prefill is compute-bound and
|
||
position-parallel, but the k iterations are serial — prefill latency
|
||
≈2.9×, typically a small fraction of end-to-end latency for G≫0.
|
||
Visible tokens: +0.
|
||
|
||
**Best-of-3 sampling.** FLOPs: with shared prompt prefill (favorable
|
||
accounting), cP + 3·cG ≈ cP + 3cG; without sharing 3c(P+G). For MBPP
|
||
(P≈150–300, G≈150–220), the *extra* FLOPs vs direct (≈2cG) are of the same
|
||
order as the loop's extra (≈1.94cP) — hence "≈matched". Wall-clock: 3×G
|
||
serial bandwidth-bound decode steps (or 3 parallel decode streams at 3×
|
||
memory); strictly worse latency than the loop unless parallelized.
|
||
Visible tokens: ≈3× (two discarded candidates). Requires a verifier or
|
||
selector to pick among samples for the overall win we report (we use
|
||
any-pass, an upper bound — see §3.4 caveat).
|
||
|
||
**Budget-CoT (50-token plan).** FLOPs: ≈c(P+G+50) plus the plan tokens'
|
||
KV in context for the remainder — the *cheapest* arm in FLOPs. Wall-clock:
|
||
+50 serial decode steps before answer tokens start (worst first-token
|
||
latency). Visible tokens: +50.
|
||
|
||
Summary: no single scalar makes these three arms "equal"; the loop
|
||
dominates on tokens and decode latency, budget-CoT on FLOPs, best-of-3 on
|
||
overall accuracy. §3.4's "≈matched FLOPs" refers to the extra-FLOPs
|
||
order-of-magnitude equivalence above, not exact equality; the honest
|
||
statement is the three-way trade-off, and we report all three axes.
|