Files
jspace/results-loop/PROTOCOL_UNIFIED.md
T
2026-07-15 12:54:58 +02:00

272 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Pre-registered protocol: unified adapter eval (written before any test numbers)
Date: 2026-07-13, after val@499, before step-799 completion. No test-set
number for the unified adapter exists at time of writing.
1. **Primary loop depth: k=2, for both tasks.** Chosen on val CE with the
tie-breaking rule: prefer the SMALLEST k whose hard-cell val CE is within
0.01 nats of the best k. (Current val: MBPP hard k2k4 = 0.004, GSM hard
k2k4 = 0.006 → both ties → k=2.) The full k-curve is secondary/descriptive.
2. **Checkpoint selection criterion (scalar, fixed now):** mean of the two
hard-cell val CEs at k=2, tasks weighted equally:
crit = (gsm_hard_k2 + mbpp_hard_k2)/2. Lowest crit among saved checkpoints
wins. (At writing: step 499, crit = (0.475+0.205)/2 = 0.340.)
3. **Primary endpoints:** (a) MBPP test pass@1 hard-bucket at k=2 vs k=0;
(b) GSM8K test accuracy hard-bucket at k=2 vs k=0. McNemar, paired by item.
Overall accuracy is secondary (known to be underpowered at n=250/256).
4. **Same-harness rule:** all k, INCLUDING k=0 baselines, measured by the
prompt-only fast-path scripts (`generate_frozen_prompt`; k=0 = plain
cached generate inside the same function). No numbers carried over from
the full-position-loop harness.
5. **Known missing control (not covered by this run):** a same-size,
no-recurrence adapter (h -> h + MLP(h) at the L13->L14 boundary, no loop,
no band re-run) trained on identical data/objective. Until it exists,
"the loop does the work (vs. 1.6M new weights anywhere doing it)" is NOT
established. Queued as the next training run. Note the k=0 column is
gated off by construction and is a sanity check only — it is not this
control.
6. **Symmetric interference check (missing):** dedicated GSM8K prompt-only
adapter as the reference for "unified costs GSM nothing". Queued. Until
then the no-interference claim is one-directional (MBPP side only).
7. **Band-location ablation (pre-registered 2026-07-13, before any arm ran).**
Arms, all else identical (adapter size/init, data, curriculum, k, scripts;
MBPP): early L2-12, mid-narrow L17-27, late L24-34 (width-matched, 11
layers); shifted L6-22 (width-matched to the original 17). Reference:
workspace L14-30 (already run, 3 seeds). Prediction: workspace-centered
arms (L14-30, L17-27) exceed early/late on hard-bucket pass@1 at k=2-4 by
a wide margin; shifted intermediate. Falsification: near-parity across
arms demotes the lens claim from "locates where to loop" to "convenient
discovery tool"; to be reported either way. Primary readout: hard-bucket
pass@1 at k=4, e400 checkpoints throughout.
8. **Language commitments for the writeup:** the k0->k1 CE collapse (e.g.
4.36->0.18) is format/template learning expected from any trained adapter
and must not be quoted as evidence of routing/planning; informative
comparisons are within k>=1 cells only. Depth ordering k2 vs k4 deltas
(0.001-0.006 nats) are inside checkpoint jitter and must be described as
"k>=2 fits hard items equally well; k=1 slightly worse."
9. **Anchor/entrance sweep (pre-registered 2026-07-14 ~03:00, before any arm
ran).** Arms: bands (13,30), (12,30), (11,30) — injection point shifted
up from L14 at fixed tap L30; plus tap-23 = (14,23). All E2B/MBPP, same
recipe, e400, primary readout hard-bucket pass@1 at k=4. Competing
predictions: (a) "L14 special" (last full-attention KV-computing layer,
lens boundary) → anchor-13 drops; (b) "KV-channel count" (anchors 11-13
add 1-3 extra KV-recomputing attention channels) → holds or improves.
Body-length confound noted: earlier anchors lengthen the loop body; if
results shift, run matched-length control (12,28) before interpreting.
10. **L9 discriminator arm (pre-registered 2026-07-14 ~10:00, before running).**
Band (9,30): anchor at L9 — the only other full-attention, KV-computing
layer below the boundary — deep in the lens's sensor regime, tap fixed at
L30. Separates the two cliff explanations: (a) "lens boundary" predicts
catastrophic (like anchors 11-13: 25-34% overall); (b) "full-attention
KV-layer entry" predicts partial recovery (clearly above the L11-13
trend, i.e. >40% overall or hard >25%). Registered prediction: (a) —
the sensor-region content dominates; layer type does not rescue it.
Same recipe/checkpoint/eval as the anchor sweep (250 items, ks 0,2,4).
---
# Outcomes vs pre-registrations (scored 2026-07-14, after all arms completed)
1. **k=2 primary depth** — held. All primary comparisons reported at k=2;
k-curves descriptive. k≥2 plateau confirmed (k=8 gen-eval flat).
2. **Checkpoint criterion** — applied as written for the unified adapter.
Separately reported: val-CE is a poor proxy for generation accuracy;
later arms therefore pre-committed to fixed steps (e400) instead.
3. **Primary endpoints (unified adapter, k=2 vs k=0)** — (a) MBPP hard
3.6% → 28.6% (direction as predicted); (b) GSM hard 0% → 6.3%, overall
10.5% → 9.0% (no overall win — the math boundary result). Both reported.
4. **Same-harness rule** — held throughout (all final tables fast-path,
k=0 included).
5. **Missing weights control** — run: trained FF adapter = 17.9% hard,
exactly the untrained-loop level. Loop-vs-weights gap established.
6. **Symmetric interference check** — run (dedicated GSM adapter);
mixed-task training regressed both tasks; reported as negative result.
7. **Band-location ablation** — prediction CONFIRMED with a caveat:
L14-30 hard 43.6% ≫ early L2-12 (23.6%, overall destroyed 22.8%) and
shifted L6-22 (21.8%, overall 29.0%). Caveat discovered: L17-27 and
L24-34 are structurally null (KV sharing; k>0 ≡ k=0 bit-identical), so
the "mid-narrow beats late" half of the prediction was untestable at
E2B; the 12B replication (no shared KV) carries that weight instead.
8. **Language commitments** — honored in PAPER.md (k0→k1 CE collapse not
cited as planning evidence; k2-vs-k4 nats described as jitter).
9. **Anchor/entrance sweep** — prediction (a) "L14 special" CONFIRMED:
anchor-13 hard 17.9%/overall 34.4%; 12: 28.6%/30.8%; 11: 25.0%/25.2%;
monotone collapse below the boundary. Tap-23 arm died in training
(never rerun); exits 27/30/32/34 within seed noise, so exit choice is
free. Matched-length control not needed (results did not shift with
body length in the informative direction).
10. **L9 discriminator** — registered prediction (a) CONFIRMED: band
(9,30) overall 14.0-21.4%, hard ≤21.4% — catastrophic, like anchors
11-13, despite L9 being a full-attention KV-computing layer. The lens
boundary, not layer type, gates the retrofit.
11. **Recurrent-regime arm (pre-registered 2026-07-15, before training).**
Huginn-style retrofit on the frozen E2B band: RecurrentAdapter
(learned A,B init α·I/(1−α)·I + zero-init MLP), h0 = norm-scaled
noise, log-uniform random depth k∈[1,16], bptt=4, same data/steps/
checkpoint rule (e400 primary) as all merge arms. Eval ks 0,2,4,8,16,32
on the 250-item MBPP set. Competing predictions: (a) "amortization is
intrinsic to frozen-band retrofits" → performance plateaus by k≈4 at
or below the merge arm's level, no depth-monotone gain; (b) "fixed-
point behavior was an artifact of our fixed-shallow-k training"
(Huginn regime transfers) → monotone hard-bucket improvement past k=8
and reduced noise-seed sensitivity after training. Secondary readout:
path independence (two noise seeds → output agreement rate) at e400.
Known risk, stated in advance: 600 steps may be far too little for
this regime (McLeish et al. use ~50B tokens); a null here bounds the
cheap-retrofit budget only, not the regime.
12. **Parcae-constrained recurrent arm (pre-registered 2026-07-15, before
training; Prairie et al. 2026 parameterization).** Same as item 11 but
A = exp(−Δt·exp(a)) diagonal → ρ(A) < 1 by construction; init exactly
the α=0.3 merge (verified bit-equal at init). ρ(A) logged every 10
steps in BOTH arms. Theory-derived predictions, stated in advance:
(a) contraction ⇒ fixed point is a function of e ⇒ the Parcae arm
SATURATES in k (no depth-monotone gain) and its converged performance
is amortizable — if so, our deflationary result is a corollary of
ρ<1, and our observed k≈34 convergence is the geometric rate 0.3^k;
(b) the UNCONSTRAINED item-11 arm either drifts toward ρ≥1 (watch the
ρ log: divergent runs should show ρ≥1 before loss spikes) or, if it
gains monotone depth-performance, does so with ρ near 1 — the edge of
stability is where genuine iteration must live. Either outcome
formalizes "the anchor coefficient is the stability dial" as
"the anchor coefficient is the spectral radius".
13. **Per-depth adapter arm + free-ACT probe (pre-registered 2026-07-15,
before training).** (a) PerDepthAdapter: one merge adapter per
iteration (n=4, Bae-style depth-wise relaxation at the entrance;
breaks time-invariance — LTV, no fixed-point guarantee), standard
curriculum, e400, eval ks 0,2,4,8. Prediction: lands at or below the
distill/rung-2 amortization ceiling (~46% hard) because depth-indexed
weights add content, not state-evolution; exceeding it would show
per-iteration expressivity was binding and amend the deflationary
claim. Depths >4 reuse adapter 4 (stated: k=8 cell is then
fixed-point-like by construction). (b) Free-ACT probe on the standard
merge arm: record per-item convergence depth (cos>0.9995) at k=8 cap.
Predictions: accuracy unchanged vs fixed k (post-convergence no-ops);
mean k_conv ≈ 3; hard-labeled items converge SLOWER than easy ones
(adaptive compute allocates like ACT without any learned halting
parameter).
--- Outcome, item 11 (scored 2026-07-15, k=16/32 cells cancelled by
decision after k<=8): PREDICTION (a) SUBSTANTIALLY CONFIRMED, with one
twist. The unconstrained arm left contraction immediately (rho(A):
0.3 -> 3.4 by step 100, plateau ~4.5) yet trained smoothly — per-iteration
norm-matching converts magnitude explosion into directional churn, so
"rho>=1 => divergence" becomes "rho>=1 => divergence OR stationary churn"
under a norm projection. Consequences as predicted: substrate damage
(easy 98.4 -> ~69% at all k>0, far exceeding any contractive arm's tax),
val CE flat k=1..16 (stationary, not progressive), hard bucket at
merge level (35.7/39.3/42.9% at k=2/4/8 — a one-item-per-depth-doubling
crawl that at k=8 reaches what the contractive merge reaches at k=4,
never approaching the amortization ceiling from above). 4x parameters
bought nothing. Depth-monotone computation did not emerge at this budget.
14. **Tied-alpha arm (pre-registered 2026-07-15, before training).**
TiedAlphaAdapter: x = (1a)⊙e + a⊙ŝ + MLP([e;ŝ]), a = σ(â) per-dim
learned, init a=0.3 everywhere (bit-equal to MergeAdapter at step 0,
verified). B tied to (1a): convex combination keeps the LTI fixed
point on the e–ŝ segment (substrate-anchored by construction),
ρ = max(a) < 1 guaranteed, +d≈1.5K params. Standard curriculum,
s0 = band(e), e400, eval ks 0,2,4,8 on 250 items. This is the one
untested cell combining parcae's learnable decay with the merge's
anchoring. Predictions: (a) substrate fidelity preserved (easy ≈
merge's 88%, unlike both rec arms' ~70%) because anchoring, not
ρ, controls fidelity; (b) hard-bucket at merge level (no significant
gain — per-dim constant α is not where capability lives, per the
adaptive-α E2B result); (c) learned a drifts slightly DOWN from 0.3
(as in parcae). If (a) holds while rec arms failed it, the
fixed-point-location dial is causally isolated: same learnable-decay
freedom, only the tie to (1a) differs from parcae.
--- Outcome, item 12 (scored 2026-07-15): prediction (a) CONFIRMED in its
dynamics half, REFUTED in its fidelity half — and the refutation is the
finding. Dynamics: rho stayed in (0,1) throughout (0.300 -> 0.292, the
optimizer drifting MORE contractive when confined to the stable region);
loss trajectory as good as or better than the unconstrained arm at every
checkpoint (the rec arm's flight to rho~4.5 was epiphenomenal — all fit
lives in the MLP); eval saturates completely (hard 42.9/42.9/39.3/39.3/
39.3 at k=2/4/8/16/32, easy flat ~71%). Fidelity: easy items were NOT
preserved (71% vs the merge's 88.5%) despite guaranteed contraction —
substrate fidelity is controlled by fixed-point LOCATION (anchored B +
curriculum), not by rho. Conclusion: stability and fidelity are
independent dials (fig_phase.png); the Parcae constraint delivers exactly
what it promises (robust training, convergence, certified tail gradients)
and exactly nothing more. Item 14 (tied-alpha) is the causal isolation of
the fidelity dial.
15. **Fidelity factorial + capacity control + seed (pre-registered
2026-07-15 ~03:15, before any of these arms ran; overnight batch).**
The fidelity loss of both rec arms (easy 88.5 -> ~71%) confounds three
deltas from the winning merge: (i) learned B, (ii) random-depth
training instead of the difficulty->depth curriculum, (iii) noise s0.
Item 14 (tied-alpha) tests (i) with anchoring. New single-variable
cells, everything else = standard merge recipe (fixed B, band(e) s0,
curriculum, e400, eval ks 0,2,4,8 on 250 items):
a. merge+randk — only (ii) changed (log-uniform k in [1,16], bptt 4).
b. merge+noises0 — only (iii) changed.
c. merge h=2048 — capacity control for the per-depth arm (6.4M
shared vs 6.4M depth-indexed): if per-depth beats the ceiling
but h2048 does not, time-variation (not capacity) is credited;
if both do, it was capacity all along.
d. parcae seed 1 — robustness of the fidelity refutation.
Predictions: (a) and (b) each cost a few points of easy at most
(anchored fixed point dominates); neither reproduces the ~17-point
drop — the culprit is the learned/free B (with item 14 as the
positive control). h2048 stays at the ceiling (hard <=46%), fidelity
intact. parcae s1 reproduces easy ~71% within seed noise.
--- Outcome, item 13a (scored 2026-07-15): prediction CONFIRMED — per-depth
lands below/at the ceiling, never above. Detail is instructive: fidelity
preserved throughout (easy 88.5/89.3/86.9 at k=2/4/8 — anchored B), but
hard-bucket content is DEPTH-STRANDED: 17.9% at k=2 (adapters 3-4, which
hold the hard-trained content, never execute), 35.7% at k=4, 42.9% at k=8
— where depths 5-8 reuse adapter 4, i.e. the architecture reverts to
shared-map iteration and the fixed-point mechanism collects the remaining
gain. Time-variation adds a fragility (content unavailable except at its
training depth) and no capability; map-sharing is load-bearing for the
anytime-usable gain. Depth-4 adapter overfit visible in val (hard k4 CE
0.188@99 -> 0.371@599) — LTV concentrates small-pool overfitting into
single depths.
--- Outcome, item 13b (scored 2026-07-15): accuracy prediction CONFIRMED
(k=8 halt run 52.0/90.2/42.9 = plateau level); convergence predictions
REFUTED. Per-item state-cosine (thresh 0.9995, k=8 cap): k_conv
distribution 4:3, 5:57, 6:47, 7:17, never-within-8:126 — mean ~7, and NO
difficulty gradient (easy 7.01 vs hard 7.00). The earlier "bit-exact by
k~3-4" was the single dynamics-probe example, not the population: outputs
plateau by k~2-4 while the state keeps drifting at 1e-3..1e-4 cosine
scale; the fixed point is an OUTPUT-stable orbit (suffix layers + decode
wash out residual state motion), not a literal state fixed point for most
prompts. Free-ACT via state-cosine therefore yields no early exit at this
threshold, and no ACT-like difficulty allocation falls out for free —
output-level halting signals would be needed. Paper's dynamics claims
softened accordingly.
--- Outcome, item 14 (scored 2026-07-15): ALL THREE PREDICTIONS CONFIRMED.
(a) Fidelity fully preserved: easy 93.4/91.0/90.2 at k=2/4/8 (merge:
92.6/88.5; parcae with identical decay freedom but untied B: ~71%) —
the free B is causally isolated as the fidelity culprit, the anchoring
tie as the protection. (b) Hard at merge level exactly (35.7/42.9/39.3 =
merge's k-curve within noise); no gain from the freedom. (c) Learned a
essentially unmoved: mean 0.298, range [0.285, 0.310], 0/1536 dims moved
>0.05 from init — the anchor coefficient is not a useful learnable DOF;
hand-tuned 0.3 was already optimal. Recipe consequence: fixed-alpha
anchored merge is the recommended design; learnable-alpha safe but
pointless, learnable-B harmful, per-depth strands the gain.
16. **Code→GSM8K cross-task transfer (pre-registered 2026-07-15 ~14:10,
before running).** The MBPP-trained loop adapter (adapter_code, s0) and
the noise-s0 variant evaluated on GSM8K test (n=256, prompt-only loop,
same harness as eval_gsmonly). Extends the transfer-distance ladder
(HumanEval tie -> LCB trained-hurts) across tasks. Predictions:
(a) hard-bucket gain ~0 (plan content is task-local; GSM8K needs
evolving state, not static plans); (b) easy items damaged at k>0
(~93 -> 50-70%), comparable to or worse than the GSM-trained merge —
substrate damage on GSM8K is perturbation-driven and content-agnostic;
(c) overall at k>0 below k=0 (no rescue). If instead hard gains
appear (>5 points), plan-shaped content is partially task-general —
would weaken the task-local claim from LCB.