item 29 arm 1: FIRST SIGNIFICANT POSITIVE — 39.1 (p=0.045 paired), every bucket a d=1 record; ablation 8-8 p=1.0 reattributes the gain: trajectory TF is a training signal for the adapter, not an inference loop

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Nils
2026-07-17 12:24:18 +02:00
co-authored by Claude Fable 5
parent f1e6a628eb
commit add87cdd72
4 changed files with 3133 additions and 0 deletions
+18
View File
@@ -900,3 +900,21 @@ Nils's morning decision: clamp test vs pivot to the hybrid/A2 line
rehearsal becomes the internalization method. All prior caveats
(position coloring, teacher=warm-start quality) carry over.
Jobs: scripts/jobs/zzz_u_traj_d1.sh, scripts/jobs/zzz_v_traj_full.sh.
IN-FLIGHT, arm 1 (tf-only, d=1) scored ~15:40: THE FIRST
SIGNIFICANT POSITIVE OF THE PROGRAM, WITH A MECHANISM TWIST.
Vals best-in-series (easy 0.218, hard 0.347). Matched 2:0 = 39.1
(drop .25 / easy .69 / hard .433 — every bucket a d=1 record;
easy IMPROVED) — clears the pre-registered >36.6 threshold; paired
vs positional d=1: 50-31, McNemar p=0.045. BUT the ablation also
scores 39.1 (burst-on vs burst-off 8-8, p=1.0): the test-time
burst is causally INERT. Attribution: trajectory teacher-forcing
is a superior TRAINING SIGNAL for the merge adapter — the ten
transition regressions teach state-folding that pays off at every
visible-token carry step — not a working inference-time loop. The
program-long pattern holds: the loop at inference has never been
causal; the loop as training scaffold just produced the largest
adapter gain ever measured. ltf plateaued ~0.085 (from 0.116) —
the autonomous operator absorbs ~25% of the driven transitions.
Next-experiment implication (Nils to confirm): apply trajectory
TF to rung A itself (nothing deleted) — if the signal generalizes,
the 57.4 headline moves.