Files
jspace/related_work/relevant_to_us.md
T

169 lines
9.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Related work — relevance to the jspace looping project
Running notes on papers in this folder: what they show, where they overlap
with our claims, and what remains ours. Update when a new PDF lands here.
Our shorthand below: "our loop" = frozen gemma-4 base + 1.6M anchor-dominant
merge adapter at L14, band L1430, prompt-only latent planning, STaR-labeled
difficulty→depth curriculum, frozen-prompt KV trick, attribution ladder
(untrained / FF / pause / loop / explicit plan), gate probe.
---
## 2511.07384 — McLeish et al., "Teaching Pretrained Language Models to
## Think Deeper with Retrofitted Recurrence" (UMD/LLNL/Tübingen, Nov 2025)
**What they do.** Convert pretrained 1B models (TinyLlama, OLMo-2, Llama-3.2)
into depth-recurrent models: remove middle layers, split the rest into
prelude → recurrent block → coda; recurrent input is a **linear adapter over
[e; s]** (prelude output ⊕ previous iterate; s₀ = random noise). Then full
continued pretraining: ~50B tokens math-heavy data, all parameters trained,
Muon optimizer, Poisson-Lognormal recurrence sampling with a curriculum
ramping mean recurrence to 32, truncated BPTT (last 8 iterations), plus a
"healing" phase (26B tokens FineWeb-Edu) to recover from layer surgery.
Result: at matched training FLOPs the retrofit beats continued-pretraining
the non-recurrent parent on GSM8K/MATH; accuracy scales with test-time
recurrence; pretrained init beats random init by ≥950B tokens of training.
**Overlap — must cite, cannot claim as novel:**
- Retrofitting recurrence into a pretrained fixed-depth model works and
beats the non-recurrent baseline (their headline, at much larger scale).
- Merge adapter at loop entry combining embedding-side and state-side
inputs ([e; s] → linear, vs our (1−α)e + α·ŝ + MLP([e;ŝ])).
- Recurrence-depth curriculum during training.
- "Pretrain traditional, then convert" (our rung-2 vision): they did the
continued-pretraining version at 1B/50B tokens. The generic claim is theirs.
**What remains ours (Paper A positioning):**
1. **Where to loop is interpretability-derived.** They pick splits by
benchmark search and name layer choice as an open problem ("future work
could identify a more optimal method for layer choice"). We derive the
band from the J-lens workspace regime and back it causally: anchor cliff
at L14, tap invariance, band-1020 location ablation ≈ 0, KV-sharing
hazard. We answer their stated open problem.
2. **Frozen base, 1.6M-param adapter, ~600 steps on a desk machine** vs all
parameters, 50B tokens, MI300A cluster — ~5 orders of magnitude apart on
the cost curve. No layer removal → no healing phase needed, and k=0
exactly recovers the base model (they cannot say that).
3. **Prompt-only latent planning + frozen-prompt KV trick: zero decode-time
cost.** Their recurrence runs on every generated token (decode cost ×r).
4. **Attribution controls** (untrained / FF / pause / explicit plan). They
have no compute-matched token-space control.
5. **Difficulty-adaptive depth** (STaR buckets, gate probe): named in their
Discussion as unsolved future work. Our gate is a first result on it.
**Worth stealing:**
- **Muon > AdamW for recurrent training** (AdamW loss-spikes to NaN) —
directly relevant when we unfreeze the band / fade-off unfreezing.
- s₀ init: theirs is random noise, ours the actual residual state — likely
why a frozen base works for us at all; expect a reviewer question here.
- FLOPs accounting convention for recurrent models: FLOPs = (6N₁ + 2N₂)D
with N₁ = params with gradients, N₂ = forward-only (truncated BPTT).
**Where to cite:** Paper A related work (primary contrast), Paper B
(layer-choice open problem → our mechanistic answer), rung-2 planning notes.
---
## 2602.14759 — Lys et al., "Inner Loop Inference for Pretrained
## Transformers: Unlocking Latent Capabilities Without Training"
## (IMT Atlantique / Sony, Mar 2026)
**What they do.** Training-free "middle looping" of frozen off-the-shelf
models (Gemma-2 2B/9B, Llama-3-8B): re-apply a block range [s, e) R times at
inference. Key findings: (1) **naive looping systematically degrades**
(distribution shift — looped activations leave the manifold the model was
trained on, some configs collapse to chance); (2) **regularized looping
rescues it**: interpolate the looped state with cached states / the baseline
state (uniform average, moving average ĥ = η·h⁽⁰⁾ + (1−η)·h⁽ᵗ⁾,
softmax auto-alignment) → modest but consistent gains across WinoGrande,
ARC, GSM8K, HellaSwag, MMLU (likelihood-scored, mostly multiple-choice).
Full start×end layer-pair sweep heatmaps. Frames looping as "logits
refinement" — depth as iterative refinement of a shared latent state.
**Overlap — must cite:**
- Their moving-average regularization η·h⁽⁰⁾ + (1−η)·h⁽ᵗ⁾ is exactly the
anchor term of our merge, minus the trained MLP. Independent confirmation
that **anchoring to the un-looped state is the thing that makes frozen-band
looping viable** — cite as convergent evidence for the anchor-dominant
design (α=0.3).
- "Training-free looping gives modest gains on a frozen model" ≈ our
untrained-loop arm (17.9% hard vs 5.5 baseline on MBPP-hard, but = FF
control). Their whole paper lives inside one cell of our attribution table.
- Layer-pair sweeps parallel our location ablation (theirs benchmark-driven,
ours interpretability-predicted then confirmed).
**What remains ours:**
1. **A trained merge** — their gains are modest by their own description;
our trained adapter roughly doubles the untrained/FF level on hard items
(17.9 → 37.4) and the pause/plan ladder bounds what the recurrence adds.
2. **Generative evals with per-item verification** (MBPP tests, GSM answer
match, Rust compile-run) vs their likelihood-scored multiple choice —
ours measures the regime where latent planning should matter.
3. **Prompt-only looping + frozen-prompt KV trick** (they loop everything,
every position; no decode-cost story).
4. Curriculum, difficulty gating, cross-task transfer arms, KV-sharing
hazard (their Gemma-2 has no KV sharing; our E2B finding that
interventions entered ≥L15 are structurally null is a hazard their
sweep methodology would silently hit on models that do share).
**Worth stealing:**
- Their distribution-shift framing of *why* naive looping fails is a clean
citable explanation for why zero-init MLP + anchor is the right
parameterization (we motivate; they demonstrate the failure mode).
- Softmax auto-alignment interpolation: a training-free adaptive α — cheap
ablation candidate against our fixed α=0.3 (relevant to the 12B α
miscalibration result).
**Where to cite:** Paper A (untrained/anchor cell; convergent evidence for
anchoring), Paper B (distribution-shift account of loop instability;
KV-sharing hazard contrast).
---
## Combined positioning takeaway
The two papers bracket us: McLeish et al. = full-retraining recurrence at
scale (expensive end), Lys et al. = zero-training looping (free end). Our
niche is the middle, and it is still open: **interpretability-chosen band +
tiny trained merge on a frozen base + prompt-only latent planning with zero
decode cost + a full attribution ladder + difficulty gating.** Neither paper
touches any of the last three items; both strengthen the premise that band
looping is a real mechanism rather than a curiosity.
## One mechanism, three training regimes (anchoring)
All three loop-entry designs are the same idea — **re-ground the looped
state in the baseline forward pass** — differing only in how much training
supports it:
| regime | anchor form | training | outcome |
|---|---|---|---|
| Lys naive | none (α_state = 1) | zero | collapse (off-manifold drift) |
| Lys regularized | fixed inference-time interpolation η·h⁽⁰⁾ + (1−η)·h⁽ᵗ⁾ | zero | modest, consistent gains |
| **ours (Paper A)** | fixed mix (1−α)e + α·ŝ + trained zero-init MLP([e;ŝ]) | 1.6M params, frozen base | large gains at 2B-effective |
| McLeish | e concatenated every iteration through a **learned** adapter | all params, 50B tokens | gains at scale; anchor constraint absorbed by retraining |
Reading: how much off-manifold state the frozen band tolerates is bounded,
and anchor strength must make up for whatever training does not provide.
Lys's naive-looping collapse is the zero-training limit of the phenomenon;
McLeish's full retraining is the opposite limit, where the manifold itself
moves to accommodate the loop and explicit anchoring dissolves into learned
injection. Our **12B α-miscalibration** (α=0.3 tuned at 2B; untrained loop
craters 72.6→23.0 at 12B) is then not a quirk but the same stability law
seen mid-spectrum: the tolerable loop share depends on the model, and a
coefficient tuned on one substrate overdrives another. This connects
directly to Lys's distribution-shift account and predicts the queued
α∈{0.1, 0.15} 12B arm should recover much of the loss; their softmax
auto-alignment (adaptive η, training-free) is the natural fallback if no
fixed α transfers. Worth a paragraph in Paper A (design justification +
12B analysis) and Paper B (stability mechanism).
**Outcome (2026-07-14, prediction confirmed):** the `12b_adaptive` MBPP
arm (adaptive anchoring, 2×/8× H100 fleet) recovers the substrate —
k=0/2/4 = 72.4/70.8/69.4 overall vs 72.6/54.6/43.0 under fixed α=0.3,
with hard monotone 11.4→18.2→27.3. The tolerable-loop-share account
called this shape in advance: stability restored by weakening the loop
share, hard gains preserved and k-monotone, residual easy erosion
(98.9→92.2) to be handled by the gate.