284 lines
18 KiB
Markdown
284 lines
18 KiB
Markdown
# Pre-registered protocol: unified adapter eval (written before any test numbers)
|
||
|
||
Date: 2026-07-13, after val@499, before step-799 completion. No test-set
|
||
number for the unified adapter exists at time of writing.
|
||
|
||
1. **Primary loop depth: k=2, for both tasks.** Chosen on val CE with the
|
||
tie-breaking rule: prefer the SMALLEST k whose hard-cell val CE is within
|
||
0.01 nats of the best k. (Current val: MBPP hard k2−k4 = 0.004, GSM hard
|
||
k2−k4 = 0.006 → both ties → k=2.) The full k-curve is secondary/descriptive.
|
||
2. **Checkpoint selection criterion (scalar, fixed now):** mean of the two
|
||
hard-cell val CEs at k=2, tasks weighted equally:
|
||
crit = (gsm_hard_k2 + mbpp_hard_k2)/2. Lowest crit among saved checkpoints
|
||
wins. (At writing: step 499, crit = (0.475+0.205)/2 = 0.340.)
|
||
3. **Primary endpoints:** (a) MBPP test pass@1 hard-bucket at k=2 vs k=0;
|
||
(b) GSM8K test accuracy hard-bucket at k=2 vs k=0. McNemar, paired by item.
|
||
Overall accuracy is secondary (known to be underpowered at n=250/256).
|
||
4. **Same-harness rule:** all k, INCLUDING k=0 baselines, measured by the
|
||
prompt-only fast-path scripts (`generate_frozen_prompt`; k=0 = plain
|
||
cached generate inside the same function). No numbers carried over from
|
||
the full-position-loop harness.
|
||
5. **Known missing control (not covered by this run):** a same-size,
|
||
no-recurrence adapter (h -> h + MLP(h) at the L13->L14 boundary, no loop,
|
||
no band re-run) trained on identical data/objective. Until it exists,
|
||
"the loop does the work (vs. 1.6M new weights anywhere doing it)" is NOT
|
||
established. Queued as the next training run. Note the k=0 column is
|
||
gated off by construction and is a sanity check only — it is not this
|
||
control.
|
||
6. **Symmetric interference check (missing):** dedicated GSM8K prompt-only
|
||
adapter as the reference for "unified costs GSM nothing". Queued. Until
|
||
then the no-interference claim is one-directional (MBPP side only).
|
||
7. **Band-location ablation (pre-registered 2026-07-13, before any arm ran).**
|
||
Arms, all else identical (adapter size/init, data, curriculum, k, scripts;
|
||
MBPP): early L2-12, mid-narrow L17-27, late L24-34 (width-matched, 11
|
||
layers); shifted L6-22 (width-matched to the original 17). Reference:
|
||
workspace L14-30 (already run, 3 seeds). Prediction: workspace-centered
|
||
arms (L14-30, L17-27) exceed early/late on hard-bucket pass@1 at k=2-4 by
|
||
a wide margin; shifted intermediate. Falsification: near-parity across
|
||
arms demotes the lens claim from "locates where to loop" to "convenient
|
||
discovery tool"; to be reported either way. Primary readout: hard-bucket
|
||
pass@1 at k=4, e400 checkpoints throughout.
|
||
8. **Language commitments for the writeup:** the k0->k1 CE collapse (e.g.
|
||
4.36->0.18) is format/template learning expected from any trained adapter
|
||
and must not be quoted as evidence of routing/planning; informative
|
||
comparisons are within k>=1 cells only. Depth ordering k2 vs k4 deltas
|
||
(0.001-0.006 nats) are inside checkpoint jitter and must be described as
|
||
"k>=2 fits hard items equally well; k=1 slightly worse."
|
||
|
||
9. **Anchor/entrance sweep (pre-registered 2026-07-14 ~03:00, before any arm
|
||
ran).** Arms: bands (13,30), (12,30), (11,30) — injection point shifted
|
||
up from L14 at fixed tap L30; plus tap-23 = (14,23). All E2B/MBPP, same
|
||
recipe, e400, primary readout hard-bucket pass@1 at k=4. Competing
|
||
predictions: (a) "L14 special" (last full-attention KV-computing layer,
|
||
lens boundary) → anchor-13 drops; (b) "KV-channel count" (anchors 11-13
|
||
add 1-3 extra KV-recomputing attention channels) → holds or improves.
|
||
Body-length confound noted: earlier anchors lengthen the loop body; if
|
||
results shift, run matched-length control (12,28) before interpreting.
|
||
|
||
10. **L9 discriminator arm (pre-registered 2026-07-14 ~10:00, before running).**
|
||
Band (9,30): anchor at L9 — the only other full-attention, KV-computing
|
||
layer below the boundary — deep in the lens's sensor regime, tap fixed at
|
||
L30. Separates the two cliff explanations: (a) "lens boundary" predicts
|
||
catastrophic (like anchors 11-13: 25-34% overall); (b) "full-attention
|
||
KV-layer entry" predicts partial recovery (clearly above the L11-13
|
||
trend, i.e. >40% overall or hard >25%). Registered prediction: (a) —
|
||
the sensor-region content dominates; layer type does not rescue it.
|
||
Same recipe/checkpoint/eval as the anchor sweep (250 items, ks 0,2,4).
|
||
|
||
---
|
||
|
||
# Outcomes vs pre-registrations (scored 2026-07-14, after all arms completed)
|
||
|
||
1. **k=2 primary depth** — held. All primary comparisons reported at k=2;
|
||
k-curves descriptive. k≥2 plateau confirmed (k=8 gen-eval flat).
|
||
2. **Checkpoint criterion** — applied as written for the unified adapter.
|
||
Separately reported: val-CE is a poor proxy for generation accuracy;
|
||
later arms therefore pre-committed to fixed steps (e400) instead.
|
||
3. **Primary endpoints (unified adapter, k=2 vs k=0)** — (a) MBPP hard
|
||
3.6% → 28.6% (direction as predicted); (b) GSM hard 0% → 6.3%, overall
|
||
10.5% → 9.0% (no overall win — the math boundary result). Both reported.
|
||
4. **Same-harness rule** — held throughout (all final tables fast-path,
|
||
k=0 included).
|
||
5. **Missing weights control** — run: trained FF adapter = 17.9% hard,
|
||
exactly the untrained-loop level. Loop-vs-weights gap established.
|
||
6. **Symmetric interference check** — run (dedicated GSM adapter);
|
||
mixed-task training regressed both tasks; reported as negative result.
|
||
7. **Band-location ablation** — prediction CONFIRMED with a caveat:
|
||
L14-30 hard 43.6% ≫ early L2-12 (23.6%, overall destroyed 22.8%) and
|
||
shifted L6-22 (21.8%, overall 29.0%). Caveat discovered: L17-27 and
|
||
L24-34 are structurally null (KV sharing; k>0 ≡ k=0 bit-identical), so
|
||
the "mid-narrow beats late" half of the prediction was untestable at
|
||
E2B; the 12B replication (no shared KV) carries that weight instead.
|
||
8. **Language commitments** — honored in PAPER.md (k0→k1 CE collapse not
|
||
cited as planning evidence; k2-vs-k4 nats described as jitter).
|
||
9. **Anchor/entrance sweep** — prediction (a) "L14 special" CONFIRMED:
|
||
anchor-13 hard 17.9%/overall 34.4%; 12: 28.6%/30.8%; 11: 25.0%/25.2%;
|
||
monotone collapse below the boundary. Tap-23 arm died in training
|
||
(never rerun); exits 27/30/32/34 within seed noise, so exit choice is
|
||
free. Matched-length control not needed (results did not shift with
|
||
body length in the informative direction).
|
||
10. **L9 discriminator** — registered prediction (a) CONFIRMED: band
|
||
(9,30) overall 14.0-21.4%, hard ≤21.4% — catastrophic, like anchors
|
||
11-13, despite L9 being a full-attention KV-computing layer. The lens
|
||
boundary, not layer type, gates the retrofit.
|
||
|
||
11. **Recurrent-regime arm (pre-registered 2026-07-15, before training).**
|
||
Huginn-style retrofit on the frozen E2B band: RecurrentAdapter
|
||
(learned A,B init α·I/(1−α)·I + zero-init MLP), h0 = norm-scaled
|
||
noise, log-uniform random depth k∈[1,16], bptt=4, same data/steps/
|
||
checkpoint rule (e400 primary) as all merge arms. Eval ks 0,2,4,8,16,32
|
||
on the 250-item MBPP set. Competing predictions: (a) "amortization is
|
||
intrinsic to frozen-band retrofits" → performance plateaus by k≈4 at
|
||
or below the merge arm's level, no depth-monotone gain; (b) "fixed-
|
||
point behavior was an artifact of our fixed-shallow-k training"
|
||
(Huginn regime transfers) → monotone hard-bucket improvement past k=8
|
||
and reduced noise-seed sensitivity after training. Secondary readout:
|
||
path independence (two noise seeds → output agreement rate) at e400.
|
||
Known risk, stated in advance: 600 steps may be far too little for
|
||
this regime (McLeish et al. use ~50B tokens); a null here bounds the
|
||
cheap-retrofit budget only, not the regime.
|
||
|
||
12. **Parcae-constrained recurrent arm (pre-registered 2026-07-15, before
|
||
training; Prairie et al. 2026 parameterization).** Same as item 11 but
|
||
A = exp(−Δt·exp(a)) diagonal → ρ(A) < 1 by construction; init exactly
|
||
the α=0.3 merge (verified bit-equal at init). ρ(A) logged every 10
|
||
steps in BOTH arms. Theory-derived predictions, stated in advance:
|
||
(a) contraction ⇒ fixed point is a function of e ⇒ the Parcae arm
|
||
SATURATES in k (no depth-monotone gain) and its converged performance
|
||
is amortizable — if so, our deflationary result is a corollary of
|
||
ρ<1, and our observed k≈3–4 convergence is the geometric rate 0.3^k;
|
||
(b) the UNCONSTRAINED item-11 arm either drifts toward ρ≥1 (watch the
|
||
ρ log: divergent runs should show ρ≥1 before loss spikes) or, if it
|
||
gains monotone depth-performance, does so with ρ near 1 — the edge of
|
||
stability is where genuine iteration must live. Either outcome
|
||
formalizes "the anchor coefficient is the stability dial" as
|
||
"the anchor coefficient is the spectral radius".
|
||
|
||
13. **Per-depth adapter arm + free-ACT probe (pre-registered 2026-07-15,
|
||
before training).** (a) PerDepthAdapter: one merge adapter per
|
||
iteration (n=4, Bae-style depth-wise relaxation at the entrance;
|
||
breaks time-invariance — LTV, no fixed-point guarantee), standard
|
||
curriculum, e400, eval ks 0,2,4,8. Prediction: lands at or below the
|
||
distill/rung-2 amortization ceiling (~46% hard) because depth-indexed
|
||
weights add content, not state-evolution; exceeding it would show
|
||
per-iteration expressivity was binding and amend the deflationary
|
||
claim. Depths >4 reuse adapter 4 (stated: k=8 cell is then
|
||
fixed-point-like by construction). (b) Free-ACT probe on the standard
|
||
merge arm: record per-item convergence depth (cos>0.9995) at k=8 cap.
|
||
Predictions: accuracy unchanged vs fixed k (post-convergence no-ops);
|
||
mean k_conv ≈ 3; hard-labeled items converge SLOWER than easy ones
|
||
(adaptive compute allocates like ACT without any learned halting
|
||
parameter).
|
||
|
||
--- Outcome, item 11 (scored 2026-07-15, k=16/32 cells cancelled by
|
||
decision after k<=8): PREDICTION (a) SUBSTANTIALLY CONFIRMED, with one
|
||
twist. The unconstrained arm left contraction immediately (rho(A):
|
||
0.3 -> 3.4 by step 100, plateau ~4.5) yet trained smoothly — per-iteration
|
||
norm-matching converts magnitude explosion into directional churn, so
|
||
"rho>=1 => divergence" becomes "rho>=1 => divergence OR stationary churn"
|
||
under a norm projection. Consequences as predicted: substrate damage
|
||
(easy 98.4 -> ~69% at all k>0, far exceeding any contractive arm's tax),
|
||
val CE flat k=1..16 (stationary, not progressive), hard bucket at
|
||
merge level (35.7/39.3/42.9% at k=2/4/8 — a one-item-per-depth-doubling
|
||
crawl that at k=8 reaches what the contractive merge reaches at k=4,
|
||
never approaching the amortization ceiling from above). 4x parameters
|
||
bought nothing. Depth-monotone computation did not emerge at this budget.
|
||
|
||
14. **Tied-alpha arm (pre-registered 2026-07-15, before training).**
|
||
TiedAlphaAdapter: x = (1−a)⊙e + a⊙ŝ + MLP([e;ŝ]), a = σ(â) per-dim
|
||
learned, init a=0.3 everywhere (bit-equal to MergeAdapter at step 0,
|
||
verified). B tied to (1−a): convex combination keeps the LTI fixed
|
||
point on the e–ŝ segment (substrate-anchored by construction),
|
||
ρ = max(a) < 1 guaranteed, +d≈1.5K params. Standard curriculum,
|
||
s0 = band(e), e400, eval ks 0,2,4,8 on 250 items. This is the one
|
||
untested cell combining parcae's learnable decay with the merge's
|
||
anchoring. Predictions: (a) substrate fidelity preserved (easy ≈
|
||
merge's 88%, unlike both rec arms' ~70%) because anchoring, not
|
||
ρ, controls fidelity; (b) hard-bucket at merge level (no significant
|
||
gain — per-dim constant α is not where capability lives, per the
|
||
adaptive-α E2B result); (c) learned a drifts slightly DOWN from 0.3
|
||
(as in parcae). If (a) holds while rec arms failed it, the
|
||
fixed-point-location dial is causally isolated: same learnable-decay
|
||
freedom, only the tie to (1−a) differs from parcae.
|
||
|
||
--- Outcome, item 12 (scored 2026-07-15): prediction (a) CONFIRMED in its
|
||
dynamics half, REFUTED in its fidelity half — and the refutation is the
|
||
finding. Dynamics: rho stayed in (0,1) throughout (0.300 -> 0.292, the
|
||
optimizer drifting MORE contractive when confined to the stable region);
|
||
loss trajectory as good as or better than the unconstrained arm at every
|
||
checkpoint (the rec arm's flight to rho~4.5 was epiphenomenal — all fit
|
||
lives in the MLP); eval saturates completely (hard 42.9/42.9/39.3/39.3/
|
||
39.3 at k=2/4/8/16/32, easy flat ~71%). Fidelity: easy items were NOT
|
||
preserved (71% vs the merge's 88.5%) despite guaranteed contraction —
|
||
substrate fidelity is controlled by fixed-point LOCATION (anchored B +
|
||
curriculum), not by rho. Conclusion: stability and fidelity are
|
||
independent dials (fig_phase.png); the Parcae constraint delivers exactly
|
||
what it promises (robust training, convergence, certified tail gradients)
|
||
and exactly nothing more. Item 14 (tied-alpha) is the causal isolation of
|
||
the fidelity dial.
|
||
|
||
15. **Fidelity factorial + capacity control + seed (pre-registered
|
||
2026-07-15 ~03:15, before any of these arms ran; overnight batch).**
|
||
The fidelity loss of both rec arms (easy 88.5 -> ~71%) confounds three
|
||
deltas from the winning merge: (i) learned B, (ii) random-depth
|
||
training instead of the difficulty->depth curriculum, (iii) noise s0.
|
||
Item 14 (tied-alpha) tests (i) with anchoring. New single-variable
|
||
cells, everything else = standard merge recipe (fixed B, band(e) s0,
|
||
curriculum, e400, eval ks 0,2,4,8 on 250 items):
|
||
a. merge+randk — only (ii) changed (log-uniform k in [1,16], bptt 4).
|
||
b. merge+noises0 — only (iii) changed.
|
||
c. merge h=2048 — capacity control for the per-depth arm (6.4M
|
||
shared vs 6.4M depth-indexed): if per-depth beats the ceiling
|
||
but h2048 does not, time-variation (not capacity) is credited;
|
||
if both do, it was capacity all along.
|
||
d. parcae seed 1 — robustness of the fidelity refutation.
|
||
Predictions: (a) and (b) each cost a few points of easy at most
|
||
(anchored fixed point dominates); neither reproduces the ~17-point
|
||
drop — the culprit is the learned/free B (with item 14 as the
|
||
positive control). h2048 stays at the ceiling (hard <=46%), fidelity
|
||
intact. parcae s1 reproduces easy ~71% within seed noise.
|
||
|
||
--- Outcome, item 13a (scored 2026-07-15): prediction CONFIRMED — per-depth
|
||
lands below/at the ceiling, never above. Detail is instructive: fidelity
|
||
preserved throughout (easy 88.5/89.3/86.9 at k=2/4/8 — anchored B), but
|
||
hard-bucket content is DEPTH-STRANDED: 17.9% at k=2 (adapters 3-4, which
|
||
hold the hard-trained content, never execute), 35.7% at k=4, 42.9% at k=8
|
||
— where depths 5-8 reuse adapter 4, i.e. the architecture reverts to
|
||
shared-map iteration and the fixed-point mechanism collects the remaining
|
||
gain. Time-variation adds a fragility (content unavailable except at its
|
||
training depth) and no capability; map-sharing is load-bearing for the
|
||
anytime-usable gain. Depth-4 adapter overfit visible in val (hard k4 CE
|
||
0.188@99 -> 0.371@599) — LTV concentrates small-pool overfitting into
|
||
single depths.
|
||
|
||
--- Outcome, item 13b (scored 2026-07-15): accuracy prediction CONFIRMED
|
||
(k=8 halt run 52.0/90.2/42.9 = plateau level); convergence predictions
|
||
REFUTED. Per-item state-cosine (thresh 0.9995, k=8 cap): k_conv
|
||
distribution 4:3, 5:57, 6:47, 7:17, never-within-8:126 — mean ~7, and NO
|
||
difficulty gradient (easy 7.01 vs hard 7.00). The earlier "bit-exact by
|
||
k~3-4" was the single dynamics-probe example, not the population: outputs
|
||
plateau by k~2-4 while the state keeps drifting at 1e-3..1e-4 cosine
|
||
scale; the fixed point is an OUTPUT-stable orbit (suffix layers + decode
|
||
wash out residual state motion), not a literal state fixed point for most
|
||
prompts. Free-ACT via state-cosine therefore yields no early exit at this
|
||
threshold, and no ACT-like difficulty allocation falls out for free —
|
||
output-level halting signals would be needed. Paper's dynamics claims
|
||
softened accordingly.
|
||
|
||
--- Outcome, item 14 (scored 2026-07-15): ALL THREE PREDICTIONS CONFIRMED.
|
||
(a) Fidelity fully preserved: easy 93.4/91.0/90.2 at k=2/4/8 (merge:
|
||
92.6/88.5; parcae with identical decay freedom but untied B: ~71%) —
|
||
the free B is causally isolated as the fidelity culprit, the anchoring
|
||
tie as the protection. (b) Hard at merge level exactly (35.7/42.9/39.3 =
|
||
merge's k-curve within noise); no gain from the freedom. (c) Learned a
|
||
essentially unmoved: mean 0.298, range [0.285, 0.310], 0/1536 dims moved
|
||
>0.05 from init — the anchor coefficient is not a useful learnable DOF;
|
||
hand-tuned 0.3 was already optimal. Recipe consequence: fixed-alpha
|
||
anchored merge is the recommended design; learnable-alpha safe but
|
||
pointless, learnable-B harmful, per-depth strands the gain.
|
||
|
||
16. **Code→GSM8K cross-task transfer (pre-registered 2026-07-15 ~14:10,
|
||
before running).** The MBPP-trained loop adapter (adapter_code, s0) and
|
||
the noise-s0 variant evaluated on GSM8K test (n=256, prompt-only loop,
|
||
same harness as eval_gsmonly). Extends the transfer-distance ladder
|
||
(HumanEval tie -> LCB trained-hurts) across tasks. Predictions:
|
||
(a) hard-bucket gain ~0 (plan content is task-local; GSM8K needs
|
||
evolving state, not static plans); (b) easy items damaged at k>0
|
||
(~93 -> 50-70%), comparable to or worse than the GSM-trained merge —
|
||
substrate damage on GSM8K is perturbation-driven and content-agnostic;
|
||
(c) overall at k>0 below k=0 (no rescue). If instead hard gains
|
||
appear (>5 points), plan-shaped content is partially task-general —
|
||
would weaken the task-local claim from LCB.
|
||
|
||
--- Amendment to item 15 (2026-07-15 ~13:15): noise-s0 arm EXCEEDED
|
||
prediction (b) upward: hard 50.0/53.6/50.0 at k=2/4/8 with easy 88-90%
|
||
— nominally the best hard cells of the project (merge best 46.4; seed
|
||
mean 37.5±5.5). Paired vs tied-alpha (only same-day per-item baseline):
|
||
discordants 5-1/3-0/3-0 in noise-s0's favor, each k p≈0.22-0.25 at n=28
|
||
— consistent direction, not individually significant. Denoising
|
||
interpretation: training the loop to reach the fixed point from noise
|
||
regularizes the content. SEED ARMS QUEUED (s1, s2, same recipe/eval,
|
||
pre-registered here): if seed-mean hard(k=4) > 46.4 (the merge's best
|
||
single cell), the recommended recipe gains noise-s0; if seed mean falls
|
||
back into 37-46, it was a lucky seed.
|