integrate LCB dissociation (untrained merge wins far-transfer, p=0.001 vs trained loop), capability panel (no MC damage), L23 exit completes sweep

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Nils
2026-07-14 19:44:07 +02:00
co-authored by Claude Fable 5
parent caef95f10f
commit 1d4cffd6ad
11 changed files with 6938 additions and 12 deletions
+56 -11
View File
@@ -34,8 +34,12 @@ degrades them), width rivals depth (trained pause registers reach 36.4%),
and verifier-assisted (oracle) sampling wins overall accuracy at matched
compute — though the *deployable* selector loses that edge entirely. What
survives is precise: the implant owns exactly the plan-dependent slice at
zero visible tokens and zero decode cost, transfers with the substrate
rather than the task, and its placement is dictated by the lens. At 12B a
zero visible tokens and zero decode cost, and its placement is dictated by
the lens. Transfer dissociates by distance: near-distribution the trained
and untrained implants tie (HumanEval); far from it (LiveCodeBench) the
*untrained* merge significantly helps while the trained content
significantly hurts — the learned content is task-local, the recurrence
substrate is general. At 12B a
constant merge coefficient destroys the substrate; a state-dependent
coefficient (3.8K parameters) restores MBPP but not Blocksworld or GSM8K —
the anchor coefficient is the stability dial that unifies this work with
@@ -174,8 +178,9 @@ hard 43.6%, overall 53.6%. L13: hard 17.9%, overall 34.4%. L9L12: overall
verified bit-identical) — the 12B model has no shared-KV layers, making it
the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34):
hard 39.346.4%, within seed spread. The lens boundary is necessary; the
exit is a free parameter. (The L23-exit arm died in training; a rerun is in
progress — the four completed exits bracket it. [L23 PENDING])
exit is a free parameter — the completed five-point exit sweep
(L23/27/30/32/34) spans hard 39.346.4% with L23 at the top (46.4% at k=2),
all within seed spread.
### 3.2 The attribution ladder
@@ -297,7 +302,28 @@ significant** (paired McNemar at k=2, 9 vs 7 discordant, p=0.80). What
transfers significantly is the *merge perturbation itself*, not the
MBPP-trained content — the cleanest evidence that off-distribution value is
substrate-shaped rather than task-memorized. (The transferred pause adapter
reaches 66.5%, hard 38.9%, consistent with the same reading.) **Rust/MultiPL-E** (Python-trained, different
reaches 66.5%, hard 38.9%, consistent with the same reading.)
**LiveCodeBench sharpens this into a dissociation** (150 newest stdin
problems, Nov 2024Apr 2025, execution-verified; no LCB training anywhere
in the pipeline; base 18.7%):
| arm (MBPP-trained where trained) | overall | hard (n=25) | vs base, paired |
|---|---|---|---|
| **untrained merge, k=4** | **24.0%** | **36.0%** | **+**, p=0.039 |
| trained loop, k=4 | 15.3% | 8.0% | , p=0.23 |
| distill FF, k=1 | 12.7% | 16.0% | ****, p=0.049 |
Far from distribution, the *trained content is a liability* (distill
significantly hurts; untrained-vs-trained-loop is 141 discordant,
p=0.001) while the *untrained anchored recurrence significantly helps*
the training-free regime of Lys et al. is the right choice off-distribution,
and the amortized-content reading of §3.3 predicts exactly this: what the
adapter learned is MBPP-shaped plan content, valuable where plans look like
MBPP plans and harmful where they don't. Transfer ordering by distance:
HumanEval (near) — trained ≈ untrained; Rust (mid) — trained helps the hard
bucket; LCB (far) — untrained wins outright. Caveats: single seed per arm,
hard n=25, one benchmark at the far end. **Rust/MultiPL-E** (Python-trained, different
language, compile-run-verified): hard 8.0%→24.0% (p=0.125 at n=25 —
directionally consistent, underpowered). **Blocksworld** MBPP-transfer:
hard 0→14.3% (task-trained: 43%). Content transfers where the substrate's
@@ -316,9 +342,22 @@ removing the easy-item perturbation tax. k=0 is the exact base model by
construction — the implant is removable at token granularity.
**General-capability panel** (ARC-Challenge, WinoGrande, HellaSwag, MMLU;
length-normalized MC scoring with the loop applied to the context span) is
running on the Spark; results will quantify what k>0 does to off-task
abilities. [PENDING — fill on completion.]
800 items each, length-normalized MC likelihood via the chat template, loop
applied to the context span). The safety answer is clean — **k>0 does not
damage general abilities**:
| arm | ARC-C | WinoGrande | HellaSwag | MMLU |
|---|---|---|---|---|
| base (k=0) | 36.0 | 55.9 | 52.3 | 30.1 |
| loop k=2 (MBPP adapter) | 36.1 | 55.3 | 49.6 | 31.3 |
| distill FF (MBPP) | 41.8 | 56.6 | 57.0 | 31.8 |
The loop arm is flat within noise (largest move 2.6 on HellaSwag,
unpaired n=800). The distill adapter *nominally improves* every benchmark
(+5.8 ARC, +4.8 HellaSwag) — consistent with §3.7's finding that these
implants carry a generically useful perturbation component, though
MC-likelihood scoring and generation quality are different regimes (see
the LCB result below before reading this as free capability).
### 3.9 Negative results with content
@@ -374,9 +413,15 @@ are small (n=55/38/25); within-ladder orderings are not individually
significant, and only the pooled hard effect and the HumanEval overall gain
survive multiple-comparison scrutiny. Bucket membership derives from greedy
labeling runs (consensus-k0 robustness check moves numbers <2 points, but
both checks share the base model; an independent difficulty proxy is an
open external check). A third architecture family was not run;
LiveCodeBench (contamination-safe) was not run; rung-2 was not run at 12B. The easy-item perturbation tax persists wherever the
both checks share the base model; an independent 12B-relabeling proxy is
running). A third architecture family was not run; rung-2 was not run at
12B. LiveCodeBench: single seed per arm, hard n=25, stdin-judged problems
only, and its newest shard (Apr 2025) is *newer than MBPP by years* but
not provably past the base model's undisclosed training cutoff — we claim
recency, not proven non-contamination. The capability panel is
MC-likelihood, not generation; its "no damage" answer does not extend to
generation quality off-distribution (LCB shows trained arms *do* hurt
there). The easy-item perturbation tax persists wherever the
gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both
arms share contamination, and memorized items land in the easy bucket, but
bucket composition is contamination-sensitive. The capability panel