revision per review: seed means in headlines, net accounting vs untrained merge, bucket definition up front, GSM12B n.s. (p=0.86), law->regularity, norm spec, compute appendix, public repo URL
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,8 +1,9 @@
|
|||||||
# Latent Planning by Workspace Recurrence: an Interpretability-Placed Implant, and What It Actually Buys
|
# Latent Planning by Workspace Recurrence: an Interpretability-Placed Implant, and What It Actually Buys
|
||||||
|
|
||||||
*Final-data draft, 2026-07-14. Base models: google/gemma-4-E2B-it and
|
*Revision draft, 2026-07-14. Base models: google/gemma-4-E2B-it and
|
||||||
gemma-4-12B-it, both frozen. Hardware: DGX Spark + rented 2×/8×H100 nodes.
|
gemma-4-12B-it, both frozen. Hardware: DGX Spark + rented 2×/8×H100 nodes.
|
||||||
Code, per-item logs, and pre-registrations: `~/jspace` (git). Statistics:
|
Code, per-item logs, and pre-registrations:
|
||||||
|
https://git.draic.info/nils/jspace (public). Statistics:
|
||||||
`results-loop/STATS.md`.*
|
`results-loop/STATS.md`.*
|
||||||
|
|
||||||
## Abstract
|
## Abstract
|
||||||
@@ -15,53 +16,56 @@ anchor-dominant merge adapter (0.03% of parameters) at the band entrance
|
|||||||
turns the non-self-map band into a stable fixed-point iteration, trained with
|
turns the non-self-map band into a stable fixed-point iteration, trained with
|
||||||
self-generated, verifier-filtered supervision. Looping the workspace over the
|
self-generated, verifier-filtered supervision. Looping the workspace over the
|
||||||
prompt ("latent planning") raises pass@1 on plan-dependent MBPP problems from
|
prompt ("latent planning") raises pass@1 on plan-dependent MBPP problems from
|
||||||
5.5% to 43.6% (seed mean 37.5±5.5), with zero visible tokens and zero
|
5.5% to 37.5±5.5 over five seeds (best seed 43.6%), with zero visible tokens
|
||||||
additional decode cost. The effect is real and highly reliable — pooled
|
and zero additional decode cost. The effect is real and highly reliable —
|
||||||
across MBPP, HumanEval, and Rust/MultiPL-E, the plan-dependent bucket moves
|
pooled across MBPP, HumanEval, and Rust/MultiPL-E, the plan-dependent bucket
|
||||||
from 4.2% to 35.6% (McNemar p≈1.5e-10) — and **placement is a law, not a
|
moves from 4.2% to 35.6% (McNemar p≈1.5e-10) — and **placement is decisive,
|
||||||
convenience**: the gain appears only when the loop enters at the
|
not a convenience**: the gain appears only when the loop enters at the
|
||||||
lens-identified boundary (L14), collapsing at L13 and below and structurally
|
lens-identified boundary (L14), collapsing at L13 and below and structurally
|
||||||
nulling above.
|
nulling above. Roughly a third of the gross effect is generic perturbation
|
||||||
|
(the untrained merge alone reaches ~18%; the bucket conditions on k=0
|
||||||
|
failure, so regression-to-mean contributes to any intervention); the
|
||||||
|
loop-specific net effect is ~+20 points over that floor.
|
||||||
|
|
||||||
But a complete attribution program deflates the mechanism's mystique.
|
A complete attribution program then deflates the mechanism's mystique: the
|
||||||
(i) **The loop's content is amortizable**: distilling the model's own
|
loop's content is *amortizable* (plan-distillation into the same adapter,
|
||||||
explicit plans into the same-size adapter — no recurrence at inference —
|
recurrence-free, matches it; nothing stacks; looping distilled weights
|
||||||
matches or exceeds the loop on the same bucket (mean over 8 runs 45.7±4.6 vs
|
degrades them), width rivals depth (trained pause registers reach 36.4%),
|
||||||
37.5±5.5, paired difference n.s.), and the two do not stack; running the
|
and compute-matched token baselines win overall accuracy outright. What
|
||||||
loop on top of the distilled adapter *degrades* it. (ii) **Width rivals
|
survives is precise: the implant owns exactly the plan-dependent slice at
|
||||||
depth**: 16 trained pause registers reach 36.4% on the same bucket.
|
zero visible tokens and zero decode cost, transfers with the substrate
|
||||||
(iii) **Compute-matched token baselines are uncomfortable**: best-of-3
|
|
||||||
sampling beats every latent arm on overall accuracy (57.2% vs ≤55.2%), and a
|
|
||||||
50-token visible plan matches the loop on the hard bucket (40.0%). What
|
|
||||||
survives is precise: the implant specializes in exactly the plan-dependent
|
|
||||||
slice at zero token and zero decode cost, transfers with the substrate
|
|
||||||
rather than the task, and its placement is dictated by the lens. At 12B a
|
rather than the task, and its placement is dictated by the lens. At 12B a
|
||||||
constant merge coefficient destroys the substrate; making the coefficient
|
constant merge coefficient destroys the substrate; a state-dependent
|
||||||
state-dependent (a 3.8K-parameter gate) restores it on MBPP
|
coefficient (3.8K parameters) restores MBPP but not Blocksworld or GSM8K —
|
||||||
(hard 11.4%→27.3% with overall preserved) but not on Blocksworld or GSM8K —
|
the anchor coefficient is the stability dial that unifies this work with
|
||||||
the stability dial that unifies this work with McLeish et al. (2511.07384)
|
McLeish et al. (2511.07384) and Lys et al. (2602.14759), and it is task-
|
||||||
and Lys et al. (2602.14759) is task- and scale-dependent.
|
and scale-dependent. Details and exact numbers: §1 and §3.
|
||||||
|
|
||||||
## 1. What this paper claims
|
## 1. What this paper claims
|
||||||
|
|
||||||
1. **A placement law.** The retrofit works if and only if the recurrence
|
*(One model family, two scales: we state findings as empirical regularities,
|
||||||
enters at the lens boundary. Entrances at L9–L13 (same adapter, data,
|
not laws.)*
|
||||||
curriculum) destroy overall accuracy (14–34% vs 52%) while recovering at
|
|
||||||
most half the hard-bucket gain; entrance at L14 preserves overall and
|
1. **A placement regularity.** The retrofit works if and only if the
|
||||||
maximizes the gain (fig_placement). Entrances at L17/L24 are *structurally
|
recurrence enters at the lens boundary. Entrances at L9–L13 (same
|
||||||
null* in this architecture: KV-sharing makes layers ≥15 reuse keys/values
|
adapter, data, curriculum) destroy overall accuracy (14–34% vs 52%)
|
||||||
computed at ≤14, so k>0 is bit-identical to k=0 — a hazard for any
|
while recovering at most half the hard-bucket gain; entrance at L14
|
||||||
retrofit method that skips the mechanistic check. Exit-layer choice is
|
preserves overall and maximizes the gain (fig_placement). Entrances at
|
||||||
nearly free (taps 27/30/32/34 within seed noise: hard 39–46%). This
|
L17/L24 are *structurally null* in this architecture: KV-sharing makes
|
||||||
answers the open "where to loop" problem named by McLeish et al., and it
|
layers ≥15 reuse keys/values computed at ≤14, so k>0 is bit-identical to
|
||||||
is causal, not correlational: the L9-entrance discriminator arm was
|
k=0 — a hazard for any retrofit method that skips the mechanistic check.
|
||||||
trained identically and fails.
|
Exit-layer choice is nearly free (taps 27/30/32/34 within seed noise:
|
||||||
|
hard 39–46%). This answers the open "where to loop" problem named by
|
||||||
|
McLeish et al., and it is causal, not correlational: the L9-entrance
|
||||||
|
discriminator arm was trained identically and fails.
|
||||||
|
|
||||||
2. **A verified, statistically solid capability gain on a narrow slice.**
|
2. **A verified, statistically solid capability gain on a narrow slice.**
|
||||||
Plan-dependent items (the model solves them with an explicit written plan
|
Plan-dependent items (the model solves them with an explicit written plan
|
||||||
but not directly): pooled across three benchmarks, 4.2%→35.6%,
|
but not directly): seed-mean 37.5±5.5 on MBPP (best 43.6%); pooled across
|
||||||
p≈1.5e-10. Overall accuracy is statistically unchanged on MBPP
|
three benchmarks, 4.2%→35.6%, p≈1.5e-10. Overall accuracy is
|
||||||
(p=0.34) and improved on HumanEval transfer (58.5%→66.5%, p=0.011).
|
statistically unchanged on MBPP (p=0.34) and improved on HumanEval
|
||||||
|
transfer (58.5%→66.5%, p=0.011). Net of the untrained-merge floor
|
||||||
|
(~18%), the loop-specific effect is ~+20 points.
|
||||||
|
|
||||||
3. **A deflationary mechanism finding.** The trained loop converges to a
|
3. **A deflationary mechanism finding.** The trained loop converges to a
|
||||||
fixed point by k≈3–4 and behaves as *amortized plan content*, not
|
fixed point by k≈3–4 and behaves as *amortized plan content*, not
|
||||||
@@ -73,7 +77,7 @@ and Lys et al. (2602.14759) is task- and scale-dependent.
|
|||||||
adapter find plan-shaped content — content that can equally be put there
|
adapter find plan-shaped content — content that can equally be put there
|
||||||
by distillation if plans are available.
|
by distillation if plans are available.
|
||||||
|
|
||||||
4. **A width-vs-depth law.** Trained pause registers (width) capture most of
|
4. **A width-vs-depth pattern.** Trained pause registers (width) capture most of
|
||||||
the plan effect on code; recurrence (depth) is needed only where a state
|
the plan effect on code; recurrence (depth) is needed only where a state
|
||||||
must *evolve* — on GSM8K generation-side carry beats registers, and on
|
must *evolve* — on GSM8K generation-side carry beats registers, and on
|
||||||
Blocksworld (pure planning, no world knowledge) the loop lifts hard-split
|
Blocksworld (pure planning, no world knowledge) the loop lifts hard-split
|
||||||
@@ -82,8 +86,10 @@ and Lys et al. (2602.14759) is task- and scale-dependent.
|
|||||||
|
|
||||||
5. **Honest economics.** The implant's costs: ≈2.9× prompt-processing FLOPs
|
5. **Honest economics.** The implant's costs: ≈2.9× prompt-processing FLOPs
|
||||||
(parallel, prefill-shaped), zero decode overhead, bit-exact KV-cache
|
(parallel, prefill-shaped), zero decode overhead, bit-exact KV-cache
|
||||||
write-in, k=0 recovers the base model exactly. Its competition at matched
|
write-in, k=0 recovers the base model exactly. Its competition at
|
||||||
FLOPs: best-of-3 sampling wins overall accuracy outright (57.2%); a
|
comparable compute (accounting in Appendix A — FLOPs, wall-clock, and
|
||||||
|
token budget do not rank the arms the same way): best-of-3 sampling wins
|
||||||
|
overall accuracy outright (57.2%); a
|
||||||
50-token visible plan ties the hard bucket. The value proposition is
|
50-token visible plan ties the hard bucket. The value proposition is
|
||||||
*only*: no visible tokens, no decode latency, and the hard-slice
|
*only*: no visible tokens, no decode latency, and the hard-slice
|
||||||
specialization (distill's 46% > budget-CoT's 40% > best-of-3's 33%).
|
specialization (distill's 46% > budget-CoT's 40% > best-of-3's 33%).
|
||||||
@@ -93,8 +99,8 @@ and Lys et al. (2602.14759) is task- and scale-dependent.
|
|||||||
damage is present *before* adapter training (untrained-loop arm) and is
|
damage is present *before* adapter training (untrained-loop arm) and is
|
||||||
not fixed by retuning α or LR. A per-position learned coefficient
|
not fixed by retuning α or LR. A per-position learned coefficient
|
||||||
α=σ(w·[e;ŝ]+b) restores MBPP (overall 69.4%, hard 11.4%→27.3%) — but
|
α=σ(w·[e;ŝ]+b) restores MBPP (overall 69.4%, hard 11.4%→27.3%) — but
|
||||||
fails to rescue Blocksworld-12B and yields only a marginal GSM8K-12B
|
fails to rescue Blocksworld-12B and yields only a nominally positive,
|
||||||
overall gain (35.9%→36.7% at k=1), the project's only overall 12B win.
|
not significant GSM8K-12B overall delta (35.9%→36.7% at k=1, p=0.86).
|
||||||
|
|
||||||
## 2. Method
|
## 2. Method
|
||||||
|
|
||||||
@@ -107,7 +113,9 @@ follow from what the readout tracks. On gemma-4-E2B: workspace ≈ L14–30 of
|
|||||||
(out-space ≠ in-space). With e = L13's output (fixed anchor) and s the
|
(out-space ≠ in-space). With e = L13's output (fixed anchor) and s the
|
||||||
fed-back, norm-matched band output:
|
fed-back, norm-matched band output:
|
||||||
|
|
||||||
L14-in = (1−α)·e + α·ŝ + MLP([e ; ŝ]), ŝ = s·|e|/|s|
|
L14-in = (1−α)·e + α·ŝ + MLP([e ; ŝ]), ŝ = s·‖e‖₂/‖s‖₂
|
||||||
|
|
||||||
|
(per-position L2 norms over the hidden dimension, computed in fp32)
|
||||||
|
|
||||||
α=0.3 constant at 2B; at 12B, α=σ(w·[e;ŝ]+b) per position (zero-init so
|
α=0.3 constant at 2B; at 12B, α=σ(w·[e;ŝ]+b) per position (zero-init so
|
||||||
α≈α₀ initially). MLP output zero-init: the untrained adapter is exactly the
|
α≈α₀ initially). MLP output zero-init: the untrained adapter is exactly the
|
||||||
@@ -136,11 +144,23 @@ Cost at k=4: ≈2.9× prefill FLOPs, **zero** decode overhead.
|
|||||||
Statistics throughout: Wilson 95% CIs; paired comparisons by exact McNemar;
|
Statistics throughout: Wilson 95% CIs; paired comparisons by exact McNemar;
|
||||||
all headline arms evaluated on the full 500-item MBPP test split (hard
|
all headline arms evaluated on the full 500-item MBPP test split (hard
|
||||||
bucket n=55), HumanEval n=164 (hard n=38), Rust/MultiPL-E n=154 (hard n=25),
|
bucket n=55), HumanEval n=164 (hard n=38), Rust/MultiPL-E n=154 (hard n=25),
|
||||||
execution-verified. Label robustness: redefining "hard" as
|
execution-verified.
|
||||||
labeled-hard ∧ k=0-fails-in-all-five-seeds (52/55 items) moves headline
|
|
||||||
numbers <2 points.
|
|
||||||
|
|
||||||
### 3.1 The placement law
|
**Bucket definition, stated up front.** Hard labels come from labeling runs
|
||||||
|
of the frozen base model on the test items themselves (direct vs
|
||||||
|
plan-in-context, greedy). This is legitimate for *descriptive* slicing but
|
||||||
|
would be circular for selection — so no arm, hyperparameter, checkpoint, or
|
||||||
|
loop depth was ever chosen using bucket results (pre-registered;
|
||||||
|
`PROTOCOL_UNIFIED.md` items 1–2, 8). Because the bucket conditions on k=0
|
||||||
|
failure, regression-to-mean inflates *any* intervention's bucket score: the
|
||||||
|
untrained merge already reaches ~18%, and we therefore report the
|
||||||
|
loop-specific effect **net of that floor** wherever attribution is claimed.
|
||||||
|
Robustness: redefining "hard" as labeled-hard ∧ k=0-fails-in-all-five-seeds
|
||||||
|
(52/55 items) moves headline numbers <2 points; both definitions share the
|
||||||
|
base model, which an independent difficulty proxy would not — we flag this
|
||||||
|
as an open external check.
|
||||||
|
|
||||||
|
### 3.1 The placement regularity
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
@@ -150,8 +170,8 @@ hard 43.6%, overall 53.6%. L13: hard 17.9%, overall 34.4%. L9–L12: overall
|
|||||||
verified bit-identical) — the 12B model has no shared-KV layers, making it
|
verified bit-identical) — the 12B model has no shared-KV layers, making it
|
||||||
the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34):
|
the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34):
|
||||||
hard 39.3–46.4%, within seed spread. The lens boundary is necessary; the
|
hard 39.3–46.4%, within seed spread. The lens boundary is necessary; the
|
||||||
exit is a free parameter. (The L23-exit arm died in training and was not
|
exit is a free parameter. (The L23-exit arm died in training; a rerun is in
|
||||||
rerun; the four completed exits bracket it.)
|
progress — the four completed exits bracket it. [L23 PENDING])
|
||||||
|
|
||||||
### 3.2 The attribution ladder
|
### 3.2 The attribution ladder
|
||||||
|
|
||||||
@@ -178,6 +198,14 @@ vs stack) is **not significant** at n=55; loop vs base *overall* is not
|
|||||||
significant on MBPP (p=0.34). The ladder's shape is reliable; its fine
|
significant on MBPP (p=0.34). The ladder's shape is reliable; its fine
|
||||||
ordering is not.
|
ordering is not.
|
||||||
|
|
||||||
|
**Net accounting.** The attribution-critical comparison is trained-loop vs
|
||||||
|
*untrained merge*, not vs base: gross 5.5→37.5 (seed mean), of which the
|
||||||
|
untrained perturbation floor is ~18 points — the loop-specific net effect
|
||||||
|
is ~+20 points. The untrained-loop and trained-FF control rows above are
|
||||||
|
from the 250-item era (hard n=28); full-bucket (n=55) reruns of both
|
||||||
|
controls, enabling the paired loop-vs-untrained test, are running and will
|
||||||
|
replace these rows. [CONTROLS-N55 PENDING]
|
||||||
|
|
||||||
### 3.3 The decisive tests: nothing stacks
|
### 3.3 The decisive tests: nothing stacks
|
||||||
|
|
||||||
If the loop performed genuine iterative computation, plan-distilled content
|
If the loop performed genuine iterative computation, plan-distilled content
|
||||||
@@ -199,17 +227,21 @@ distinctive value is that it finds this content *without* plan supervision
|
|||||||
|
|
||||||
### 3.4 Compute-matched honesty
|
### 3.4 Compute-matched honesty
|
||||||
|
|
||||||
At approximately matched FLOPs, token-space baselines are strong: best-of-3
|
At approximately matched FLOPs (Appendix A gives the accounting, separated
|
||||||
sampling wins overall accuracy against every latent arm (57.2%,
|
into FLOPs, wall-clock, and token budget), token-space baselines are strong:
|
||||||
|
best-of-3 sampling wins overall accuracy against every latent arm (57.2%,
|
||||||
CI [52.8, 61.5], vs loop 53.6 [49.2, 57.9] — point estimate higher, CIs
|
CI [52.8, 61.5], vs loop 53.6 [49.2, 57.9] — point estimate higher, CIs
|
||||||
overlap) by preserving easy items perfectly while sampling rescues some hard
|
overlap) by preserving easy items perfectly while sampling rescues some hard
|
||||||
ones. A 50-token visible plan ties the loop's hard bucket. The latent
|
ones. A 50-token visible plan ties the loop's hard bucket. Both baselines
|
||||||
implant's surviving advantages are qualitative: zero visible tokens (silent),
|
are being rerun with per-item logs to enable paired tests against the latent
|
||||||
zero decode overhead (prefill-parallel; sampling and CoT pay serially at
|
arms; until those land, the overall-accuracy comparison rests on overlapping
|
||||||
bandwidth-bound decode), and the hard-slice crown under distillation (46% vs
|
CIs and is stated as point-estimate-level. [BASELINES-PI PENDING]
|
||||||
40% budget-CoT vs 33% best-of-3). For deployment this means: the implant is
|
The latent implant's surviving advantages are qualitative: zero visible
|
||||||
a *latency/token-budget* technology with a side specialization in
|
tokens (silent), zero decode overhead (prefill-parallel; sampling and CoT
|
||||||
plan-dependent items — not an accuracy technology.
|
pay serially at bandwidth-bound decode), and the hard-slice crown under
|
||||||
|
distillation (46% vs 40% budget-CoT vs 33% best-of-3). For deployment this
|
||||||
|
means: the implant is a *latency/token-budget* technology with a side
|
||||||
|
specialization in plan-dependent items — not an accuracy technology.
|
||||||
|
|
||||||
### 3.5 Width vs depth, and the task boundary
|
### 3.5 Width vs depth, and the task boundary
|
||||||
|
|
||||||
@@ -219,7 +251,7 @@ prompt-side result entirely (no variant beats the weights control
|
|||||||
prompt-side), but generation-side *carry* — recurrence across token steps —
|
prompt-side), but generation-side *carry* — recurrence across token steps —
|
||||||
doubles the pause control on hard items: arithmetic's serial state evolves
|
doubles the pause control on hard items: arithmetic's serial state evolves
|
||||||
during the answer. Blocksworld at 2B is the purest case: base 0% on hard
|
during the answer. Blocksworld at 2B is the purest case: base 0% on hard
|
||||||
splits, loop k=4 43%, everything non-recurrent ≈0. The law: **plans are
|
splits, loop k=4 43%, everything non-recurrent ≈0. The pattern: **plans are
|
||||||
wide; execution is deep.** Retrofit recurrence pays off precisely where a
|
wide; execution is deep.** Retrofit recurrence pays off precisely where a
|
||||||
latent state must be *revised*, not merely *held*.
|
latent state must be *revised*, not merely *held*.
|
||||||
|
|
||||||
@@ -233,9 +265,11 @@ the damage precedes adapter training; α=0.15 and LR retuning do not fix it
|
|||||||
(47.6/52.6% overall). The state-dependent coefficient does, on MBPP:
|
(47.6/52.6% overall). The state-dependent coefficient does, on MBPP:
|
||||||
overall 69.4% (base 72.4%), hard 11.4%→27.3%. It does **not** rescue
|
overall 69.4% (base 72.4%), hard 11.4%→27.3%. It does **not** rescue
|
||||||
Blocksworld-12B (easy items destroyed at k=4; constant-α had reached hard
|
Blocksworld-12B (easy items destroyed at k=4; constant-α had reached hard
|
||||||
40% but also destroyed easy) and yields only +0.8 points overall on
|
40% but also destroyed easy) and yields a **nominally positive, not
|
||||||
GSM8K-12B (35.9→36.7 at k=1, hard 1.6→10.6) — the sole overall-accuracy win
|
significant** overall delta on GSM8K-12B (35.9→36.7 at k=1; paired McNemar
|
||||||
of the program, and a marginal one. Conclusion: the anchor coefficient is
|
on 32 discordant items, p=0.86; hard 1.6→10.6) — no arm anywhere in the
|
||||||
|
program produced a statistically significant overall gain at 12B.
|
||||||
|
Conclusion: the anchor coefficient is
|
||||||
the load-bearing stability control, its correct *form* (not just value)
|
the load-bearing stability control, its correct *form* (not just value)
|
||||||
changes with scale, and per-task tuning remains unavoidable.
|
changes with scale, and per-task tuning remains unavoidable.
|
||||||
|
|
||||||
@@ -314,7 +348,7 @@ proved a genuine rival, not a strawman (§3.2, §3.5).
|
|||||||
What remains distinct here: interpretability-derived placement with causal
|
What remains distinct here: interpretability-derived placement with causal
|
||||||
validation; a fully frozen base with bit-exact k=0 and zero-decode-cost KV
|
validation; a fully frozen base with bit-exact k=0 and zero-decode-cost KV
|
||||||
write-in; the complete attribution ladder including compute-matched
|
write-in; the complete attribution ladder including compute-matched
|
||||||
token-space baselines and stacking tests; the width/depth task law; and the
|
token-space baselines and stacking tests; the width/depth task pattern; and the
|
||||||
amortizability finding itself.
|
amortizability finding itself.
|
||||||
|
|
||||||
## 5. Limitations
|
## 5. Limitations
|
||||||
@@ -324,10 +358,9 @@ are small (n=55/38/25); within-ladder orderings are not individually
|
|||||||
significant, and only the pooled hard effect and the HumanEval overall gain
|
significant, and only the pooled hard effect and the HumanEval overall gain
|
||||||
survive multiple-comparison scrutiny. Bucket membership derives from greedy
|
survive multiple-comparison scrutiny. Bucket membership derives from greedy
|
||||||
labeling runs (consensus-k0 robustness check moves numbers <2 points, but
|
labeling runs (consensus-k0 robustness check moves numbers <2 points, but
|
||||||
both checks share the base model). Best-of-3/budget-CoT lack per-item logs
|
both checks share the base model; an independent difficulty proxy is an
|
||||||
(no paired tests against them). The L23 exit arm and a third architecture
|
open external check). A third architecture family was not run;
|
||||||
family were not run; LiveCodeBench (contamination-safe) was not run; rung-2
|
LiveCodeBench (contamination-safe) was not run; rung-2 was not run at 12B. The easy-item perturbation tax persists wherever the
|
||||||
was not run at 12B. The easy-item perturbation tax persists wherever the
|
|
||||||
gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both
|
gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both
|
||||||
arms share contamination, and memorized items land in the easy bucket, but
|
arms share contamination, and memorized items land in the easy bucket, but
|
||||||
bucket composition is contamination-sensitive. The capability panel
|
bucket composition is contamination-sensitive. The capability panel
|
||||||
@@ -351,3 +384,39 @@ difficulty; keep k=0 as the exact base model. What it buys: the
|
|||||||
plan-dependent slice at zero tokens and zero decode cost. What it does not
|
plan-dependent slice at zero tokens and zero decode cost. What it does not
|
||||||
buy: overall accuracy beyond what matched-compute sampling already delivers.
|
buy: overall accuracy beyond what matched-compute sampling already delivers.
|
||||||
Both halves of that sentence are the contribution.
|
Both halves of that sentence are the contribution.
|
||||||
|
|
||||||
|
## Appendix A: compute accounting (FLOPs / wall-clock / tokens, separated)
|
||||||
|
|
||||||
|
Let P = prompt tokens, G = generated tokens, c = FLOPs per token per full
|
||||||
|
forward pass. The band is 17 of 35 decoder layers at E2B (fraction
|
||||||
|
f≈0.486) and 10 of 48 at 12B (f≈0.208).
|
||||||
|
|
||||||
|
**Latent loop, k=4, prompt-only.** Prefill: 1 base pass + 4 band passes
|
||||||
|
over prompt positions = (1+4f)·cP ≈ **2.94·cP** at E2B (1.83× at 12B —
|
||||||
|
the overhead *shrinks* with scale because lens bands grow sublinearly).
|
||||||
|
Decode: exactly cG (looped states written to the KV cache once;
|
||||||
|
bit-exactness verified). Wall-clock: prefill is compute-bound and
|
||||||
|
position-parallel, but the k iterations are serial — prefill latency
|
||||||
|
≈2.9×, typically a small fraction of end-to-end latency for G≫0.
|
||||||
|
Visible tokens: +0.
|
||||||
|
|
||||||
|
**Best-of-3 sampling.** FLOPs: with shared prompt prefill (favorable
|
||||||
|
accounting), cP + 3·cG ≈ cP + 3cG; without sharing 3c(P+G). For MBPP
|
||||||
|
(P≈150–300, G≈150–220), the *extra* FLOPs vs direct (≈2cG) are of the same
|
||||||
|
order as the loop's extra (≈1.94cP) — hence "≈matched". Wall-clock: 3×G
|
||||||
|
serial bandwidth-bound decode steps (or 3 parallel decode streams at 3×
|
||||||
|
memory); strictly worse latency than the loop unless parallelized.
|
||||||
|
Visible tokens: ≈3× (two discarded candidates). Requires a verifier or
|
||||||
|
selector to pick among samples for the overall win we report (we use
|
||||||
|
any-pass, an upper bound — see §3.4 caveat).
|
||||||
|
|
||||||
|
**Budget-CoT (50-token plan).** FLOPs: ≈c(P+G+50) plus the plan tokens'
|
||||||
|
KV in context for the remainder — the *cheapest* arm in FLOPs. Wall-clock:
|
||||||
|
+50 serial decode steps before answer tokens start (worst first-token
|
||||||
|
latency). Visible tokens: +50.
|
||||||
|
|
||||||
|
Summary: no single scalar makes these three arms "equal"; the loop
|
||||||
|
dominates on tokens and decode latency, budget-CoT on FLOPs, best-of-3 on
|
||||||
|
overall accuracy. §3.4's "≈matched FLOPs" refers to the extra-FLOPs
|
||||||
|
order-of-magnitude equivalence above, not exact equality; the honest
|
||||||
|
statement is the three-way trade-off, and we report all three axes.
|
||||||
|
|||||||
Reference in New Issue
Block a user