Files
jspace/results-loop/PROTOCOL_UNIFIED.md
T

423 lines
26 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Pre-registered protocol: unified adapter eval (written before any test numbers)
Date: 2026-07-13, after val@499, before step-799 completion. No test-set
number for the unified adapter exists at time of writing.
1. **Primary loop depth: k=2, for both tasks.** Chosen on val CE with the
tie-breaking rule: prefer the SMALLEST k whose hard-cell val CE is within
0.01 nats of the best k. (Current val: MBPP hard k2k4 = 0.004, GSM hard
k2k4 = 0.006 → both ties → k=2.) The full k-curve is secondary/descriptive.
2. **Checkpoint selection criterion (scalar, fixed now):** mean of the two
hard-cell val CEs at k=2, tasks weighted equally:
crit = (gsm_hard_k2 + mbpp_hard_k2)/2. Lowest crit among saved checkpoints
wins. (At writing: step 499, crit = (0.475+0.205)/2 = 0.340.)
3. **Primary endpoints:** (a) MBPP test pass@1 hard-bucket at k=2 vs k=0;
(b) GSM8K test accuracy hard-bucket at k=2 vs k=0. McNemar, paired by item.
Overall accuracy is secondary (known to be underpowered at n=250/256).
4. **Same-harness rule:** all k, INCLUDING k=0 baselines, measured by the
prompt-only fast-path scripts (`generate_frozen_prompt`; k=0 = plain
cached generate inside the same function). No numbers carried over from
the full-position-loop harness.
5. **Known missing control (not covered by this run):** a same-size,
no-recurrence adapter (h -> h + MLP(h) at the L13->L14 boundary, no loop,
no band re-run) trained on identical data/objective. Until it exists,
"the loop does the work (vs. 1.6M new weights anywhere doing it)" is NOT
established. Queued as the next training run. Note the k=0 column is
gated off by construction and is a sanity check only — it is not this
control.
6. **Symmetric interference check (missing):** dedicated GSM8K prompt-only
adapter as the reference for "unified costs GSM nothing". Queued. Until
then the no-interference claim is one-directional (MBPP side only).
7. **Band-location ablation (pre-registered 2026-07-13, before any arm ran).**
Arms, all else identical (adapter size/init, data, curriculum, k, scripts;
MBPP): early L2-12, mid-narrow L17-27, late L24-34 (width-matched, 11
layers); shifted L6-22 (width-matched to the original 17). Reference:
workspace L14-30 (already run, 3 seeds). Prediction: workspace-centered
arms (L14-30, L17-27) exceed early/late on hard-bucket pass@1 at k=2-4 by
a wide margin; shifted intermediate. Falsification: near-parity across
arms demotes the lens claim from "locates where to loop" to "convenient
discovery tool"; to be reported either way. Primary readout: hard-bucket
pass@1 at k=4, e400 checkpoints throughout.
8. **Language commitments for the writeup:** the k0->k1 CE collapse (e.g.
4.36->0.18) is format/template learning expected from any trained adapter
and must not be quoted as evidence of routing/planning; informative
comparisons are within k>=1 cells only. Depth ordering k2 vs k4 deltas
(0.001-0.006 nats) are inside checkpoint jitter and must be described as
"k>=2 fits hard items equally well; k=1 slightly worse."
9. **Anchor/entrance sweep (pre-registered 2026-07-14 ~03:00, before any arm
ran).** Arms: bands (13,30), (12,30), (11,30) — injection point shifted
up from L14 at fixed tap L30; plus tap-23 = (14,23). All E2B/MBPP, same
recipe, e400, primary readout hard-bucket pass@1 at k=4. Competing
predictions: (a) "L14 special" (last full-attention KV-computing layer,
lens boundary) → anchor-13 drops; (b) "KV-channel count" (anchors 11-13
add 1-3 extra KV-recomputing attention channels) → holds or improves.
Body-length confound noted: earlier anchors lengthen the loop body; if
results shift, run matched-length control (12,28) before interpreting.
10. **L9 discriminator arm (pre-registered 2026-07-14 ~10:00, before running).**
Band (9,30): anchor at L9 — the only other full-attention, KV-computing
layer below the boundary — deep in the lens's sensor regime, tap fixed at
L30. Separates the two cliff explanations: (a) "lens boundary" predicts
catastrophic (like anchors 11-13: 25-34% overall); (b) "full-attention
KV-layer entry" predicts partial recovery (clearly above the L11-13
trend, i.e. >40% overall or hard >25%). Registered prediction: (a) —
the sensor-region content dominates; layer type does not rescue it.
Same recipe/checkpoint/eval as the anchor sweep (250 items, ks 0,2,4).
---
# Outcomes vs pre-registrations (scored 2026-07-14, after all arms completed)
1. **k=2 primary depth** — held. All primary comparisons reported at k=2;
k-curves descriptive. k≥2 plateau confirmed (k=8 gen-eval flat).
2. **Checkpoint criterion** — applied as written for the unified adapter.
Separately reported: val-CE is a poor proxy for generation accuracy;
later arms therefore pre-committed to fixed steps (e400) instead.
3. **Primary endpoints (unified adapter, k=2 vs k=0)** — (a) MBPP hard
3.6% → 28.6% (direction as predicted); (b) GSM hard 0% → 6.3%, overall
10.5% → 9.0% (no overall win — the math boundary result). Both reported.
4. **Same-harness rule** — held throughout (all final tables fast-path,
k=0 included).
5. **Missing weights control** — run: trained FF adapter = 17.9% hard,
exactly the untrained-loop level. Loop-vs-weights gap established.
6. **Symmetric interference check** — run (dedicated GSM adapter);
mixed-task training regressed both tasks; reported as negative result.
7. **Band-location ablation** — prediction CONFIRMED with a caveat:
L14-30 hard 43.6% ≫ early L2-12 (23.6%, overall destroyed 22.8%) and
shifted L6-22 (21.8%, overall 29.0%). Caveat discovered: L17-27 and
L24-34 are structurally null (KV sharing; k>0 ≡ k=0 bit-identical), so
the "mid-narrow beats late" half of the prediction was untestable at
E2B; the 12B replication (no shared KV) carries that weight instead.
8. **Language commitments** — honored in PAPER.md (k0→k1 CE collapse not
cited as planning evidence; k2-vs-k4 nats described as jitter).
9. **Anchor/entrance sweep** — prediction (a) "L14 special" CONFIRMED:
anchor-13 hard 17.9%/overall 34.4%; 12: 28.6%/30.8%; 11: 25.0%/25.2%;
monotone collapse below the boundary. Tap-23 arm died in training
(never rerun); exits 27/30/32/34 within seed noise, so exit choice is
free. Matched-length control not needed (results did not shift with
body length in the informative direction).
10. **L9 discriminator** — registered prediction (a) CONFIRMED: band
(9,30) overall 14.0-21.4%, hard ≤21.4% — catastrophic, like anchors
11-13, despite L9 being a full-attention KV-computing layer. The lens
boundary, not layer type, gates the retrofit.
11. **Recurrent-regime arm (pre-registered 2026-07-15, before training).**
Huginn-style retrofit on the frozen E2B band: RecurrentAdapter
(learned A,B init α·I/(1−α)·I + zero-init MLP), h0 = norm-scaled
noise, log-uniform random depth k∈[1,16], bptt=4, same data/steps/
checkpoint rule (e400 primary) as all merge arms. Eval ks 0,2,4,8,16,32
on the 250-item MBPP set. Competing predictions: (a) "amortization is
intrinsic to frozen-band retrofits" → performance plateaus by k≈4 at
or below the merge arm's level, no depth-monotone gain; (b) "fixed-
point behavior was an artifact of our fixed-shallow-k training"
(Huginn regime transfers) → monotone hard-bucket improvement past k=8
and reduced noise-seed sensitivity after training. Secondary readout:
path independence (two noise seeds → output agreement rate) at e400.
Known risk, stated in advance: 600 steps may be far too little for
this regime (McLeish et al. use ~50B tokens); a null here bounds the
cheap-retrofit budget only, not the regime.
12. **Parcae-constrained recurrent arm (pre-registered 2026-07-15, before
training; Prairie et al. 2026 parameterization).** Same as item 11 but
A = exp(−Δt·exp(a)) diagonal → ρ(A) < 1 by construction; init exactly
the α=0.3 merge (verified bit-equal at init). ρ(A) logged every 10
steps in BOTH arms. Theory-derived predictions, stated in advance:
(a) contraction ⇒ fixed point is a function of e ⇒ the Parcae arm
SATURATES in k (no depth-monotone gain) and its converged performance
is amortizable — if so, our deflationary result is a corollary of
ρ<1, and our observed k≈34 convergence is the geometric rate 0.3^k;
(b) the UNCONSTRAINED item-11 arm either drifts toward ρ≥1 (watch the
ρ log: divergent runs should show ρ≥1 before loss spikes) or, if it
gains monotone depth-performance, does so with ρ near 1 — the edge of
stability is where genuine iteration must live. Either outcome
formalizes "the anchor coefficient is the stability dial" as
"the anchor coefficient is the spectral radius".
13. **Per-depth adapter arm + free-ACT probe (pre-registered 2026-07-15,
before training).** (a) PerDepthAdapter: one merge adapter per
iteration (n=4, Bae-style depth-wise relaxation at the entrance;
breaks time-invariance — LTV, no fixed-point guarantee), standard
curriculum, e400, eval ks 0,2,4,8. Prediction: lands at or below the
distill/rung-2 amortization ceiling (~46% hard) because depth-indexed
weights add content, not state-evolution; exceeding it would show
per-iteration expressivity was binding and amend the deflationary
claim. Depths >4 reuse adapter 4 (stated: k=8 cell is then
fixed-point-like by construction). (b) Free-ACT probe on the standard
merge arm: record per-item convergence depth (cos>0.9995) at k=8 cap.
Predictions: accuracy unchanged vs fixed k (post-convergence no-ops);
mean k_conv ≈ 3; hard-labeled items converge SLOWER than easy ones
(adaptive compute allocates like ACT without any learned halting
parameter).
--- Outcome, item 11 (scored 2026-07-15, k=16/32 cells cancelled by
decision after k<=8): PREDICTION (a) SUBSTANTIALLY CONFIRMED, with one
twist. The unconstrained arm left contraction immediately (rho(A):
0.3 -> 3.4 by step 100, plateau ~4.5) yet trained smoothly — per-iteration
norm-matching converts magnitude explosion into directional churn, so
"rho>=1 => divergence" becomes "rho>=1 => divergence OR stationary churn"
under a norm projection. Consequences as predicted: substrate damage
(easy 98.4 -> ~69% at all k>0, far exceeding any contractive arm's tax),
val CE flat k=1..16 (stationary, not progressive), hard bucket at
merge level (35.7/39.3/42.9% at k=2/4/8 — a one-item-per-depth-doubling
crawl that at k=8 reaches what the contractive merge reaches at k=4,
never approaching the amortization ceiling from above). 4x parameters
bought nothing. Depth-monotone computation did not emerge at this budget.
14. **Tied-alpha arm (pre-registered 2026-07-15, before training).**
TiedAlphaAdapter: x = (1a)⊙e + a⊙ŝ + MLP([e;ŝ]), a = σ(â) per-dim
learned, init a=0.3 everywhere (bit-equal to MergeAdapter at step 0,
verified). B tied to (1a): convex combination keeps the LTI fixed
point on the e–ŝ segment (substrate-anchored by construction),
ρ = max(a) < 1 guaranteed, +d≈1.5K params. Standard curriculum,
s0 = band(e), e400, eval ks 0,2,4,8 on 250 items. This is the one
untested cell combining parcae's learnable decay with the merge's
anchoring. Predictions: (a) substrate fidelity preserved (easy ≈
merge's 88%, unlike both rec arms' ~70%) because anchoring, not
ρ, controls fidelity; (b) hard-bucket at merge level (no significant
gain — per-dim constant α is not where capability lives, per the
adaptive-α E2B result); (c) learned a drifts slightly DOWN from 0.3
(as in parcae). If (a) holds while rec arms failed it, the
fixed-point-location dial is causally isolated: same learnable-decay
freedom, only the tie to (1a) differs from parcae.
--- Outcome, item 12 (scored 2026-07-15): prediction (a) CONFIRMED in its
dynamics half, REFUTED in its fidelity half — and the refutation is the
finding. Dynamics: rho stayed in (0,1) throughout (0.300 -> 0.292, the
optimizer drifting MORE contractive when confined to the stable region);
loss trajectory as good as or better than the unconstrained arm at every
checkpoint (the rec arm's flight to rho~4.5 was epiphenomenal — all fit
lives in the MLP); eval saturates completely (hard 42.9/42.9/39.3/39.3/
39.3 at k=2/4/8/16/32, easy flat ~71%). Fidelity: easy items were NOT
preserved (71% vs the merge's 88.5%) despite guaranteed contraction —
substrate fidelity is controlled by fixed-point LOCATION (anchored B +
curriculum), not by rho. Conclusion: stability and fidelity are
independent dials (fig_phase.png); the Parcae constraint delivers exactly
what it promises (robust training, convergence, certified tail gradients)
and exactly nothing more. Item 14 (tied-alpha) is the causal isolation of
the fidelity dial.
15. **Fidelity factorial + capacity control + seed (pre-registered
2026-07-15 ~03:15, before any of these arms ran; overnight batch).**
The fidelity loss of both rec arms (easy 88.5 -> ~71%) confounds three
deltas from the winning merge: (i) learned B, (ii) random-depth
training instead of the difficulty->depth curriculum, (iii) noise s0.
Item 14 (tied-alpha) tests (i) with anchoring. New single-variable
cells, everything else = standard merge recipe (fixed B, band(e) s0,
curriculum, e400, eval ks 0,2,4,8 on 250 items):
a. merge+randk — only (ii) changed (log-uniform k in [1,16], bptt 4).
b. merge+noises0 — only (iii) changed.
c. merge h=2048 — capacity control for the per-depth arm (6.4M
shared vs 6.4M depth-indexed): if per-depth beats the ceiling
but h2048 does not, time-variation (not capacity) is credited;
if both do, it was capacity all along.
d. parcae seed 1 — robustness of the fidelity refutation.
Predictions: (a) and (b) each cost a few points of easy at most
(anchored fixed point dominates); neither reproduces the ~17-point
drop — the culprit is the learned/free B (with item 14 as the
positive control). h2048 stays at the ceiling (hard <=46%), fidelity
intact. parcae s1 reproduces easy ~71% within seed noise.
--- Outcome, item 13a (scored 2026-07-15): prediction CONFIRMED — per-depth
lands below/at the ceiling, never above. Detail is instructive: fidelity
preserved throughout (easy 88.5/89.3/86.9 at k=2/4/8 — anchored B), but
hard-bucket content is DEPTH-STRANDED: 17.9% at k=2 (adapters 3-4, which
hold the hard-trained content, never execute), 35.7% at k=4, 42.9% at k=8
— where depths 5-8 reuse adapter 4, i.e. the architecture reverts to
shared-map iteration and the fixed-point mechanism collects the remaining
gain. Time-variation adds a fragility (content unavailable except at its
training depth) and no capability; map-sharing is load-bearing for the
anytime-usable gain. Depth-4 adapter overfit visible in val (hard k4 CE
0.188@99 -> 0.371@599) — LTV concentrates small-pool overfitting into
single depths.
--- Outcome, item 13b (scored 2026-07-15): accuracy prediction CONFIRMED
(k=8 halt run 52.0/90.2/42.9 = plateau level); convergence predictions
REFUTED. Per-item state-cosine (thresh 0.9995, k=8 cap): k_conv
distribution 4:3, 5:57, 6:47, 7:17, never-within-8:126 — mean ~7, and NO
difficulty gradient (easy 7.01 vs hard 7.00). The earlier "bit-exact by
k~3-4" was the single dynamics-probe example, not the population: outputs
plateau by k~2-4 while the state keeps drifting at 1e-3..1e-4 cosine
scale; the fixed point is an OUTPUT-stable orbit (suffix layers + decode
wash out residual state motion), not a literal state fixed point for most
prompts. Free-ACT via state-cosine therefore yields no early exit at this
threshold, and no ACT-like difficulty allocation falls out for free —
output-level halting signals would be needed. Paper's dynamics claims
softened accordingly.
--- Outcome, item 14 (scored 2026-07-15): ALL THREE PREDICTIONS CONFIRMED.
(a) Fidelity fully preserved: easy 93.4/91.0/90.2 at k=2/4/8 (merge:
92.6/88.5; parcae with identical decay freedom but untied B: ~71%) —
the free B is causally isolated as the fidelity culprit, the anchoring
tie as the protection. (b) Hard at merge level exactly (35.7/42.9/39.3 =
merge's k-curve within noise); no gain from the freedom. (c) Learned a
essentially unmoved: mean 0.298, range [0.285, 0.310], 0/1536 dims moved
>0.05 from init — the anchor coefficient is not a useful learnable DOF;
hand-tuned 0.3 was already optimal. Recipe consequence: fixed-alpha
anchored merge is the recommended design; learnable-alpha safe but
pointless, learnable-B harmful, per-depth strands the gain.
16. **Code→GSM8K cross-task transfer (pre-registered 2026-07-15 ~14:10,
before running).** The MBPP-trained loop adapter (adapter_code, s0) and
the noise-s0 variant evaluated on GSM8K test (n=256, prompt-only loop,
same harness as eval_gsmonly). Extends the transfer-distance ladder
(HumanEval tie -> LCB trained-hurts) across tasks. Predictions:
(a) hard-bucket gain ~0 (plan content is task-local; GSM8K needs
evolving state, not static plans); (b) easy items damaged at k>0
(~93 -> 50-70%), comparable to or worse than the GSM-trained merge —
substrate damage on GSM8K is perturbation-driven and content-agnostic;
(c) overall at k>0 below k=0 (no rescue). If instead hard gains
appear (>5 points), plan-shaped content is partially task-general —
would weaken the task-local claim from LCB.
--- Amendment to item 15 (2026-07-15 ~13:15): noise-s0 arm EXCEEDED
prediction (b) upward: hard 50.0/53.6/50.0 at k=2/4/8 with easy 88-90%
— nominally the best hard cells of the project (merge best 46.4; seed
mean 37.5±5.5). Paired vs tied-alpha (only same-day per-item baseline):
discordants 5-1/3-0/3-0 in noise-s0's favor, each k p≈0.22-0.25 at n=28
— consistent direction, not individually significant. Denoising
interpretation: training the loop to reach the fixed point from noise
regularizes the content. SEED ARMS QUEUED (s1, s2, same recipe/eval,
pre-registered here): if seed-mean hard(k=4) > 46.4 (the merge's best
single cell), the recommended recipe gains noise-s0; if seed mean falls
back into 37-46, it was a lucky seed.
--- Outcome, item 15c (h2048 capacity control, scored 2026-07-15): the
per-depth exoneration is CLEAN — shared 6.4M params reach hard 42.9/53.6/
50.0 at k=2/4/8 vs per-depth's 17.9/35.7/42.9 at the same capacity;
time-variation is strictly worse than weight-sharing at matched params.
Fidelity prediction confirmed and exceeded (easy 95.1% at k=2 — best
looped fidelity of the project; 90.2% at k=4/8). Ceiling prediction
(hard <= 46%) REFUTED UPWARD like noise-s0: k=4/8 at 53.6/50.0. Two
independent variations (noise s0, 4x MLP) now sit at 50-54% where the
original merge reached 46.4 — suggests 46.4 was an UNDER-estimate of the
recipe family's level, not a ceiling it defined. The distill-parity
deflation claim is unaffected statistically (53.6 vs 45.7 at hard n=28
is within noise) but the language "every regime tops out at the same
ceiling" should become "at the same level within noise" — pending the
noise-s0 seed arms.
--- Outcome, item 15d (parcae seed 1, scored 2026-07-15): CONFIRMED —
the fidelity refutation replicates. easy 70.5/73.0/72.1 at k=2/4/8
(seed 0: 72.1/71.3/70.5); hard 32.1/39.3/35.7 (seed 0: 42.9/42.9/39.3,
ordinary seed spread at n=28). Two-seed conclusion: contraction-with-
free-B loses ~17 points of easy items regardless of seed; the phase
diagram's Parcae point is solid.
--- Outcome, item 16 (code->GSM8K transfer, scored 2026-07-15): ALL THREE
PREDICTIONS CONFIRMED, emphatically. MBPP-trained loop on GSM8K: hard
0.8-1.6% at every k (prediction a: ~0 gain — plan content is task-local);
easy 93.1 -> 27.6-44.8% (prediction b: damaged, in fact WORSE than the
GSM-trained merge's 48%); overall strictly below k=0 at every k>0
(prediction c). noise-s0 variant identical (easy 34.5, hard 1.6). The
transfer-distance ladder ends cleanly: near (HumanEval) tie, far-code
(LCB) trained-hurts, cross-task (GSM8K) trained-content actively toxic
while gaining nothing. Task-locality of the learned content is now a
three-point monotone result.
--- Closure of the item-15b/15c "ceiling nudged upward" question
(2026-07-15, after ns seeds): LUCKY SEED, per the pre-registered rule.
noise-s0 hard(k=4) across seeds: 53.6 / 39.3 / 35.7 -> seed mean 42.9,
inside the 37-46 band. Fidelity across seeds intact (easy 90.2-94.3 —
the factorial conclusion is seed-robust); the 50-54% cells (ns seed 0,
h2048 single seed) were upper-tail draws of the same distribution the
merge's 46.4 came from. No recipe amendment; the abstract's original
"same level within noise" framing stands; single-cell records are not
levels — only seed means are.
17. **GSM-only, current recipe (pre-registered 2026-07-15 ~20:45, before
running).** train_merge_unified.py --tasks gsm: MergeAdapter, prompt-
only loop, curriculum, GSM8K data ONLY — removes the mixed-task
interference confound from the adapter_uni run, completing the
"winning recipe trained on GSM" question. Eval: prompt-only, n=256,
ks 0,1,2,4, e400. Predictions: (a) hard <= 10% at every k (supervision
density is structural: ~3 answer tokens; the recipe's dense-output
ingredient cannot exist here); (b) easy damaged at k>0 (to 40-70%);
(c) overall never beats k=0. If hard exceeds 15% or overall beats
k=0, task interference in the mixed run was masking a real GSM
capability — would reopen the GSM chapter.
Scope note (item 17): the design-space arms of items 11-15 are NOT
crossed with GSM8K, deliberately. Exclusion by dominance: fidelity-
failing regimes (rec, parcae) cannot improve on a task MORE fidelity-
fragile than MBPP; architecture-failing (per-depth) and equivalent
(tied-alpha -> merge) and k-placement-only (randk) and same-family
(noise-s0, h2048) variants have no mechanism by which task change
could invert their MBPP verdict. Only the recipe family's best member
(this item) is informative on GSM8K.
--- Outcome, item 17 (GSM-only, current recipe, scored 2026-07-15):
predictions (a) and (b) CONFIRMED, (c) nominally exceeded but not
meaningfully. hard 8.7/5.5/4.7% at k=1/2/4 (below the 10% bar; nowhere
near the 15% reopen threshold); easy 93.1 -> 48-52% at k>0; overall
11.7/10.9/10.2 vs k0's 10.5 — the k=1 cell is +1.2 points nominal
(~3 items at n=256, not significant), the rest below. Removing the
mixed-task interference bought ~2 points over adapter_uni (9.4 -> 11.7
at k=1) — interference was real but marginal, not masking a capability.
The GSM8K chapter is closed: the recipe family's best member, trained
on GSM alone in the correct regime, delivers no usable gain and the
standard fidelity damage; combined with the scope note, the boundary
claim (structural: supervision density + state-evolution bottleneck)
is fully supported.
18. **E1: learned per-prompt halting gate (pre-registered 2026-07-16
~00:20, before any arm runs; PLAN_SELFPACED.md).** HaltingMergeAdapter:
frozen-recipe merge + ACT-style halting head on the last prompt
position's workspace state; soft state-mixture training, CE + lambda *
E[iters], penalty warmup at step 100; NO difficulty curriculum (mixed
batches — the gate must discover the allocation). k_max=4, e400/e600
checkpoints, deploy = sequential halting at 0.5 cumulative mass,
generation via frozen-prompt at per-item k*. Arms: lambda in
{0, 1e-3, 1e-2}, seed 0. Eval: 250 items, vs anchors k=0 (0.488),
uniform merge k=4 (0.512/0.885/0.464), probe-gate E0 (0.520/0.975/0.286).
Predictions: (a) some lambda gives overall >= 0.512 at mean E[k] <=
2.4 (60% of uniform-4); (b) easy >= 0.95 at that lambda; (c) k*-vs-hard
point-biserial r > 0.3; (d) hard >= 0.286 (beats E0's frozen probe).
Collapse (E[k] pinned at 1 or 4 for all lambda) falsifies E1 and
triggers the plan's kill criterion. lambda=0 control isolates whether
the CE gradient alone moves the gate (expected: barely — penalty
provides the pressure).
Item 18 amendment (2026-07-16 ~23:45, before results): arms run on a
rented 4xH100 node in parallel instead of the Spark queue; a fourth
arm (lambda=1e-3, seed 1) is added for immediate seed replication of
the expected-winner penalty. Spark's queued gate jobs will be dropped
to avoid duplication. Everything else per registration.
--- Outcome, item 18 (scored 2026-07-16 ~00:40): predictions (b), (c)
REFUTED, (a) marginal miss, (d) trivial pass. All arms converge to
UNIFORM depth (lambda 0/1e-3/1e-2 -> E[k] 4/2-or-4/1; the two 1e-3 seeds
picked different plateaus — degenerate penalty landscape), r = 0.000
everywhere. Mechanism identified and consistent with prior findings:
teacher-forced CE is depth-flat (stationarity), so CE provides no
per-item depth gradient; the penalty alone cannot teach selectivity.
The state DOES carry the signal (E0 probe: train acc 1.0) — the failure
is the training signal, not the representation. E1-as-designed is dead;
kill criterion NOT fully triggered (E2 untested, and the mechanism
points at a repair).
19. **E1b: label-supervised halting head (pre-registered 2026-07-16
~00:45, before running).** Freeze the curriculum merge (adapter_code
s0); train ONLY the halting head (BCE): target halt=0 at iterations
below the label's depth (easy->1, hard->4, per STaR label), halt=1 at
or above it. 300 steps, mixed batches, head-only params. Eval: gated
eval as item 18, n=250. Predictions: (a) r(k*, hard) > 0.5 (the head
is a trained difficulty classifier now); (b) easy >= 95% at k*=1
(near-E0's 97.5); (c) hard >= 35.7% (>= best uniform arm, via better
recall than E0's frozen probe: more than 18/28 hard items routed
deep); (d) overall >= 52.0 at E[k] <= 2.2. If (c) fails while (a,b)
hold, halting-head recall saturates at probe level and gate quality,
not gate training, is the binding constraint.
--- Outcome, item 19 / E1b (scored 2026-07-16 ~01:15): prediction (c)
CONFIRMED (hard 39.3 >= 35.7 at mean k* 2.18), (a) FAILED at r=0.217
(selectivity real — hard routed 2x deeper than easy (2.18 vs 1.08), the
program's first nonzero gate correlation — but weak at deploy), (b,d)
FAILED for a traced design reason: halted_k_per_item lacked k*=0, so easy
items were forced through >=1 iteration and landed on the merge's WORST
easy depth (k=1: 85.2%); E0's 97.5% came precisely from k=0 routing.
E1c amendment (pre-registered before running, same session): pre-loop
halt consult on s_0 enabling k*=0; targets easy->0, hard->4; threshold
0.5 unchanged (calibration deferred unless E1c misses). Predictions:
easy >= 95%, hard >= 35.7%, r >= 0.4, overall >= 51.2 at E[k] <= 1.5.