diff --git a/results-loop/PROTOCOL_UNIFIED.md b/results-loop/PROTOCOL_UNIFIED.md index e7c15e4..d9bf387 100644 --- a/results-loop/PROTOCOL_UNIFIED.md +++ b/results-loop/PROTOCOL_UNIFIED.md @@ -64,3 +64,40 @@ number for the unified adapter exists at time of writing. trend, i.e. >40% overall or hard >25%). Registered prediction: (a) — the sensor-region content dominates; layer type does not rescue it. Same recipe/checkpoint/eval as the anchor sweep (250 items, ks 0,2,4). + +--- + +# Outcomes vs pre-registrations (scored 2026-07-14, after all arms completed) + +1. **k=2 primary depth** — held. All primary comparisons reported at k=2; + k-curves descriptive. k≥2 plateau confirmed (k=8 gen-eval flat). +2. **Checkpoint criterion** — applied as written for the unified adapter. + Separately reported: val-CE is a poor proxy for generation accuracy; + later arms therefore pre-committed to fixed steps (e400) instead. +3. **Primary endpoints (unified adapter, k=2 vs k=0)** — (a) MBPP hard + 3.6% → 28.6% (direction as predicted); (b) GSM hard 0% → 6.3%, overall + 10.5% → 9.0% (no overall win — the math boundary result). Both reported. +4. **Same-harness rule** — held throughout (all final tables fast-path, + k=0 included). +5. **Missing weights control** — run: trained FF adapter = 17.9% hard, + exactly the untrained-loop level. Loop-vs-weights gap established. +6. **Symmetric interference check** — run (dedicated GSM adapter); + mixed-task training regressed both tasks; reported as negative result. +7. **Band-location ablation** — prediction CONFIRMED with a caveat: + L14-30 hard 43.6% ≫ early L2-12 (23.6%, overall destroyed 22.8%) and + shifted L6-22 (21.8%, overall 29.0%). Caveat discovered: L17-27 and + L24-34 are structurally null (KV sharing; k>0 ≡ k=0 bit-identical), so + the "mid-narrow beats late" half of the prediction was untestable at + E2B; the 12B replication (no shared KV) carries that weight instead. +8. **Language commitments** — honored in PAPER.md (k0→k1 CE collapse not + cited as planning evidence; k2-vs-k4 nats described as jitter). +9. **Anchor/entrance sweep** — prediction (a) "L14 special" CONFIRMED: + anchor-13 hard 17.9%/overall 34.4%; 12: 28.6%/30.8%; 11: 25.0%/25.2%; + monotone collapse below the boundary. Tap-23 arm died in training + (never rerun); exits 27/30/32/34 within seed noise, so exit choice is + free. Matched-length control not needed (results did not shift with + body length in the informative direction). +10. **L9 discriminator** — registered prediction (a) CONFIRMED: band + (9,30) overall 14.0-21.4%, hard ≤21.4% — catastrophic, like anchors + 11-13, despite L9 being a full-attention KV-computing layer. The lens + boundary, not layer type, gates the retrofit.