score outcomes against pre-registered protocol items 1-10
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -64,3 +64,40 @@ number for the unified adapter exists at time of writing.
|
|||||||
trend, i.e. >40% overall or hard >25%). Registered prediction: (a) —
|
trend, i.e. >40% overall or hard >25%). Registered prediction: (a) —
|
||||||
the sensor-region content dominates; layer type does not rescue it.
|
the sensor-region content dominates; layer type does not rescue it.
|
||||||
Same recipe/checkpoint/eval as the anchor sweep (250 items, ks 0,2,4).
|
Same recipe/checkpoint/eval as the anchor sweep (250 items, ks 0,2,4).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Outcomes vs pre-registrations (scored 2026-07-14, after all arms completed)
|
||||||
|
|
||||||
|
1. **k=2 primary depth** — held. All primary comparisons reported at k=2;
|
||||||
|
k-curves descriptive. k≥2 plateau confirmed (k=8 gen-eval flat).
|
||||||
|
2. **Checkpoint criterion** — applied as written for the unified adapter.
|
||||||
|
Separately reported: val-CE is a poor proxy for generation accuracy;
|
||||||
|
later arms therefore pre-committed to fixed steps (e400) instead.
|
||||||
|
3. **Primary endpoints (unified adapter, k=2 vs k=0)** — (a) MBPP hard
|
||||||
|
3.6% → 28.6% (direction as predicted); (b) GSM hard 0% → 6.3%, overall
|
||||||
|
10.5% → 9.0% (no overall win — the math boundary result). Both reported.
|
||||||
|
4. **Same-harness rule** — held throughout (all final tables fast-path,
|
||||||
|
k=0 included).
|
||||||
|
5. **Missing weights control** — run: trained FF adapter = 17.9% hard,
|
||||||
|
exactly the untrained-loop level. Loop-vs-weights gap established.
|
||||||
|
6. **Symmetric interference check** — run (dedicated GSM adapter);
|
||||||
|
mixed-task training regressed both tasks; reported as negative result.
|
||||||
|
7. **Band-location ablation** — prediction CONFIRMED with a caveat:
|
||||||
|
L14-30 hard 43.6% ≫ early L2-12 (23.6%, overall destroyed 22.8%) and
|
||||||
|
shifted L6-22 (21.8%, overall 29.0%). Caveat discovered: L17-27 and
|
||||||
|
L24-34 are structurally null (KV sharing; k>0 ≡ k=0 bit-identical), so
|
||||||
|
the "mid-narrow beats late" half of the prediction was untestable at
|
||||||
|
E2B; the 12B replication (no shared KV) carries that weight instead.
|
||||||
|
8. **Language commitments** — honored in PAPER.md (k0→k1 CE collapse not
|
||||||
|
cited as planning evidence; k2-vs-k4 nats described as jitter).
|
||||||
|
9. **Anchor/entrance sweep** — prediction (a) "L14 special" CONFIRMED:
|
||||||
|
anchor-13 hard 17.9%/overall 34.4%; 12: 28.6%/30.8%; 11: 25.0%/25.2%;
|
||||||
|
monotone collapse below the boundary. Tap-23 arm died in training
|
||||||
|
(never rerun); exits 27/30/32/34 within seed noise, so exit choice is
|
||||||
|
free. Matched-length control not needed (results did not shift with
|
||||||
|
body length in the informative direction).
|
||||||
|
10. **L9 discriminator** — registered prediction (a) CONFIRMED: band
|
||||||
|
(9,30) overall 14.0-21.4%, hard ≤21.4% — catastrophic, like anchors
|
||||||
|
11-13, despite L9 being a full-attention KV-computing layer. The lens
|
||||||
|
boundary, not layer type, gates the retrofit.
|
||||||
|
|||||||
Reference in New Issue
Block a user