Compare commits

..
113 Commits
Author SHA1 Message Date
NilsandClaude Fable 5 f1bd5daf62 MATRIX.md: item 32 row + final architecture verdict
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 02:10:40 +02:00
NilsandClaude Fable 5 21b605c13a item 32 scored: latent paper net-harmful — TF 20.3 (exposure catastrophe), ST 27.0 (quantization noise), analog 39.1 wins; loop=plan-machine / tokens=executor division final; architecture tree closed
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 02:10:02 +02:00
NilsandClaude Fable 5 f0b7942c7a item 32 pre-registered (Nils's synthesis): discrete latent chain — lens-snapped token embeddings fed back via zero-init projector (ST top-32, TF/free-running arms, frozen arm-1 merge); sym_iterate in carry_common, trainer/eval wiring
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 00:16:08 +02:00
NilsandClaude Fable 5 3db6dd8fef MATRIX.md: full item 1-31 results matrix (arcs, positives, walls, open threads); overview PDF v3 updated with items 29-31 finals
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 21:24:13 +02:00
NilsandClaude Fable 5 7ac7f9651a item 31 scored: KV-memory flat (39.5 vs 39.1, p=1.0); all gates unmoved at -10 — vanishing-gradient init confound recorded; -3 rerun is the loose thread
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 21:04:46 +02:00
NilsandClaude Fable 5 96afac769e item 30 scored: readiness strongly decodable (AUC 0.80) but NOT recurrence-specific (FF 0.79); early-stop loses everywhere, oracle only +3.9 — head is instrumentation, not a stopping lever; extension/deferral gating stays live
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 20:04:06 +02:00
NilsandClaude Fable 5 64a20750bd item 29 complete: trajectory TF = adapter training signal for hybrid regimes (39.1, p=0.045), not an internalization mechanism (job 2 answer-only 15.6 < 19.1 plateau; FR term harmful)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 14:46:54 +02:00
NilsandClaude Fable 5 b18ffbd739 item 31 pre-registered (Nils's design): synthetic memory tokens — burst states → per-band-layer KV prefix via KVMemoryAdapter (in-place attention wrap, bit-exact disarmed, gated silent init); composed with frozen item-29 arm-1; +29 tf+fr in-flight note (30.5, FR term hurts)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 13:25:25 +02:00
NilsandClaude Fable 5 11604d0949 item 30 pre-registered (Nils's design): carried state as metacognitive signal — answer-readiness head on L30 at line boundaries, carry-vs-FF AUC as primary, theta sweep + oracle stop via item-20 LUT methodology
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 12:57:05 +02:00
NilsandClaude Fable 5 c7167d512f overview PDF v2: add the prior arc (items 1-20) — retrofit discovery, band-location/anchor ablations, adapter-class factorial, transfer ladder, GSM boundary, gate program
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 12:36:38 +02:00
NilsandClaude Fable 5 3e02c6e24e INTERNALIZATION_OVERVIEW.pdf: program overview items 21-29 (architecture, ladder+controls, microscopy, lens-supervision series, trajectory TF result, conclusions & open moves)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 12:33:29 +02:00
NilsandClaude Fable 5 add87cdd72 item 29 arm 1: FIRST SIGNIFICANT POSITIVE — 39.1 (p=0.045 paired), every bucket a d=1 record; ablation 8-8 p=1.0 reattributes the gain: trajectory TF is a training signal for the adapter, not an inference loop
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 12:24:18 +02:00
NilsandClaude Fable 5 f1e6a628eb item 29 pre-registered (Nils's design): trajectory teacher-forcing — 10 waypoint transitions supervised independently (T[i-1]→T[i]), TF and TF+FR arms; job 1 d=1 step-span, job 2 answer-only full-CoT span
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 10:34:44 +02:00
NilsandClaude Fable 5 deae500ad0 item 26 scored: staging supervision actively harmful (24.2 alone, 15.6 combined) — blanket pre-'=' forcing corrupts the board's just-in-time schedule; overnight program complete, items 25-28 all scored
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 07:18:09 +02:00
NilsandClaude Fable 5 1af6cb5e45 item 28 scored: distillation succeeded (cos 0.113→0.044), function didn't follow — 30.1 matched / 29.7 ablated, easy damaged to 48.3; state-side supervision 2x2 complete and uniformly null; read-side clamp is item 29 candidate
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 03:41:43 +02:00
NilsandClaude Fable 5 67f9c3091c item 27 scored: 34.0 best-in-series with zero pauses, but paired ns (burst 10-6 p=0.45; vs positional 44-38 p=0.58) — ceiling holds; pause tape confirmed ~worthless (ablation 32.4)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 02:12:54 +02:00
NilsandClaude Fable 5 6542f80339 item 28 job: fix eval tags to ts50
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 00:46:57 +02:00
NilsandClaude Fable 5 bbf8a02c6b item 28 pre-registered (Nils's variant): teacher-state distillation — frozen full-cot teacher's band-exit state at step end, cosine into burst s^10; λ amended 1.0→5.0 pre-run (smoke: baseline cos-dist 0.113)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 00:46:19 +02:00
NilsandClaude Fable 5 62ce30e557 item 25 closed (truncated): writing is not computing — lce 1.9 with 2:12 flat at 31.2 (log-only; eval json never written due to kill); λ=1.0 arm cancelled; item 27 promoted
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 00:25:32 +02:00
NilsandClaude Fable 5 44e814e78c item 27 pre-registered: zero-pause internal band looping — single M=10 in-place burst at prompt end, iteration-aligned lens targets, ii0 inference ablation; carry_steps gains inplace updates, trainer gains --inner-iters/--base-pauses
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 00:21:20 +02:00
NilsandClaude Fable 5 0cffce876b item 26 pre-registered: result-staging lens supervision at pre-'=' positions (no causal leakage — low loss requires computation); arms lg03 and lt03+lg03; item-25 in-flight note (lce 10.3→2.3)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 23:23:13 +02:00
NilsandClaude Fable 5 64abfe66f3 item 25 pre-registered: latent process supervision via differentiable lens readout — pause j trained to lens-encode deleted-step token j (λ=0.3/1.0 arms); carry_logits gains return_states
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 23:03:56 +02:00
NilsandClaude Fable 5 f31e6e7e30 item 24 closed (stopped by Nils): trainable band doesn't improve fit either; E2-L internalization line closed on all four axes; LoopLoRA k=0 bit-exactness validated in passing
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 22:55:10 +02:00
NilsandClaude Fable 5 fe89820e64 item 24 pre-registered: d=1 retry with loop-only band-LoRA r=16 (4.8M params, whole band, k=0 bit-exact); trainer/eval gain --bandlora; smoke-tested train+eval
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 21:44:01 +02:00
NilsandClaude Fable 5 7a141948fa item 23 scored: both control axes confirm — pp30 29.3 (within ±5 of 31.6, 3x pauses bought nothing), x600 no recovery; d=1 ceiling is structural, ladder chapter closed
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 20:08:54 +02:00
NilsandClaude Fable 5 46b5e1b802 whiteboard microscopy on hard(795)/drop(430) specimens: FF slips at the 3-term sum (155), carry re-expands and defers (165 ✓); boards show live continuation-plan vs premature digit-commit at identical prefixes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 19:48:58 +02:00
NilsandClaude Fable 5 f143c67b3d probe_discount: single-prompt whiteboard microscopy (lens table + digit probs, carry vs FF control) on the 80/15% problem
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 19:14:23 +02:00
NilsandClaude Fable 5 001a4ca5f3 item 23 pre-registered: d=1 ceiling controls (x600 undertrain arm, pp30 pause-bandwidth arm, +-5pt decision rule); trainer gains --tag-suffix
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 16:27:54 +02:00
NilsandClaude Fable 5 7d60f37cf1 item 22 scored: ladder breaks at d=1 (57.4->31.6), plateaus ~19 at d=2/3 — 2x cold floor (d=3 vs base McNemar p=0.0025); latent loop trades easy reliability for hard/drop reach
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 15:29:42 +02:00
NilsandClaude Fable 5 4e744e7d94 26B-A4B lens mapped: 30 layers (not 48), same early-persist/terminal-motor family as E4B+31B; patched save delivered full regimes.pt incl. per-layer ignition; 5-scale figure complete
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 11:03:38 +02:00
NilsandClaude Fable 5 20f7014f5b item 22 pre-registered: E2-L rung B internalization ladder (front-first deletion, 10 pauses/step, warm-started d=1..3); trainer gains --drop-steps/--warm-start/--steps/--lr; smoke-tested 2 steps on Spark
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 10:13:01 +02:00
NilsandClaude Fable 5 272a7b1d1f 31B regimes parsed (early persist bump L7-16, motor only terminal — E2B signature absent, matches E4B); exp4_regimes.py takes explicit out path / defaults next to input jbar (fix pushed to node before 26B stage)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 09:54:52 +02:00
NilsandClaude Fable 5 34497b8835 item 21 McNemar scored: drop-bucket edge significant (2:6 p=0.0094, survives Bonferroni x4); overall A-vs-B ns; + session handoff, fig cleanup
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 09:49:21 +02:00
Nils 051e90e805 item 21: GSM carry-cot 57.4% (5x prior best); control isolates whiteboard to drop-bucket; E2-N/A2 planned 2026-07-16 09:40:23 +02:00
Nils 1618c206ae item 20 scored; item 21 (GSM carry-CoT + control) pre-registered and queued 2026-07-16 02:42:46 +02:00
NilsandClaude Fable 5 4b9e168838 item 20: gate threshold curve + oracle bound (probs pass, merge per-item LUT, offline sweep)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 02:03:12 +02:00
NilsandClaude Fable 5 74f04d124f node_bootstrap.sh v2: one-command node setup encoding every paid-for lesson (workspace-only, snapshot pinning, credential-aware sync, contract workers, graceful stop) + gpuq_presign for credential-less nodes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 01:51:44 +02:00
NilsandClaude Fable 5 21908a0e33 E1c scored: easy routing solved (95.9% at 0.11 iters), hard recall regressed; threshold calibration = identified next knob
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 01:48:14 +02:00
NilsandClaude Fable 5 a643b4a146 E4B regime artifacts (registry row, diagnostics, scan log, updated figure)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 01:36:57 +02:00
NilsandClaude Fable 5 332f7f3622 E4B regime map (ckpt-salvaged, 150 prompts): no clear workspace signature — sensor L11-22, motor from L23, persistence bump absent; candidate band inside KV-shared zone
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 01:35:38 +02:00
NilsandClaude Fable 5 dec27a4017 gpuq worker: STOP sentinel for graceful drain (never kill mid-job again)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 01:21:00 +02:00
NilsandClaude Fable 5 550d5cbe56 gpuq data contract: declarative per-job in/out sync (pull inputs, push outputs every 120s + at exit), gpuq_sync.sh helper
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 01:17:33 +02:00
NilsandClaude Fable 5 684e793eef E1c: k*=0 routing + pre-loop halt target (item 19 scored, E1c pre-registered)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 23:51:44 +02:00
NilsandClaude Fable 5 f2a48ce338 E1b trainer: label-supervised halting head on frozen merge (item 19)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 23:43:37 +02:00
NilsandClaude Fable 5 dafbd8fd25 item 18 scored (uniform-depth collapse, CE depth-flatness mechanism); item 19 E1b pre-registered
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 23:42:31 +02:00
NilsandClaude Fable 5 ba95d2aa3d item 18 amendment: E1 arms moved to 4xH100 parallel + seed-1 arm added (pre-results)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 23:28:38 +02:00
NilsandClaude Fable 5 357a5add06 E1 halting-gate machinery (adapter, trainer, gated eval) + pre-registration item 18
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 23:00:39 +02:00
NilsandClaude Fable 5 5f58bc8aed PLAN_SELFPACED.md: v2 prototype plan for learned workspace-compute gating (E0 numbers grounded, E1-E3 pre-registered)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 22:50:35 +02:00
NilsandClaude Fable 5 600f965247 regimes figure: sensor/workspace/motor shading restored (rule-derived, matches original hand-shaded bands)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 22:09:09 +02:00
NilsandClaude Fable 5 dc43769118 fig_regimes.py: reproducible cross-model depth-regimes figure (E2B+12B filled, 3 scans pending)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 22:05:29 +02:00
NilsandClaude Fable 5 8187b79b50 track the E2B lens-regime figures (were gitignored, disk-only)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 22:02:42 +02:00
NilsandClaude Fable 5 3a0a6a96ae INDEX.md (task-oriented repo guide), REGIMES.json -> results/, phase diagram updated with full sweep (anchored family, seed means)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 21:52:26 +02:00
NilsandClaude Fable 5 5d681a57b5 REGIMES: 31B structure (60L, d=5376, no KV-sharing key found)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 21:48:42 +02:00
NilsandClaude Fable 5 8ebb326745 REGIMES.json: canonical per-model lens registry (bands, KV boundaries, pinned revisions)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 21:45:12 +02:00
NilsandClaude Fable 5 cb0f1d950c item 17 scored: GSM-only current recipe fails as predicted; GSM chapter closed
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 21:14:58 +02:00
NilsandClaude Fable 5 2cd1f37f5f protocol: principled exclusion note for the design-space x GSM8K matrix
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 21:03:38 +02:00
NilsandClaude Fable 5 ed5455277b pre-register item 17: GSM-only prompt-only arm (last missing cell of the GSM question)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 21:01:23 +02:00
NilsandClaude Fable 5 c97d9d9cda kcurves final: all 10 panels filled
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 20:57:45 +02:00
NilsandClaude Fable 5 d81afef488 items 16 + 15-closure scored: cross-task transfer toxic (task-locality confirmed 3-point); noise-s0 50%+ was a lucky seed
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 20:57:44 +02:00
NilsandClaude Fable 5 1e6473e1aa parcae seed 1: fidelity refutation replicates (easy ~71-73% both seeds)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 20:56:28 +02:00
NilsandClaude Fable 5 0f386ea29d fig_kcurves.py: restore original output filename
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 18:34:18 +02:00
NilsandClaude Fable 5 b6d54c2726 LESSONS: thermal hard-freeze pattern and diagnosis by exclusion
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 18:30:01 +02:00
NilsandClaude Fable 5 b17d7faea0 rename kcurves figure -> fig_kcurves2.png (cache-bust)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 18:00:07 +02:00
NilsandClaude Fable 5 873dc351a0 kcurves updated: noise-s0 and h2048 panels filled (both cross the distill line)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 17:56:04 +02:00
NilsandClaude Fable 5 cf9e38f858 h2048 outcome: per-depth exonerated (sharing >> time-variation at matched params); ceiling nudged upward
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 15:00:32 +02:00
NilsandClaude Fable 5 713fbdce76 randk + noise-s0 eval JSONs (landed pre-crash)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 14:15:26 +02:00
NilsandClaude Fable 5 f0e524cf68 noise-s0 exceeds prediction upward; seed arms queued and pre-registered
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:01:14 +02:00
NilsandClaude Fable 5 ebf0eb8f49 pre-register item 16: code->GSM8K cross-task transfer
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 12:54:58 +02:00
NilsandClaude Fable 5 b013ca27f1 kcurves: full design-space grid (10 panels, distill reference line, auto-fills pending arms)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 12:42:29 +02:00
NilsandClaude Fable 5 4455d4bb9d ladder figure: base vs FF vs loops vs distill (fig_loop_vs_ff)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 12:26:22 +02:00
NilsandClaude Fable 5 13643fa504 abstract rewritten: fold in regime sweep (two independent dials, B-tie isolation, ceiling survival, output-orbit)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 11:42:15 +02:00
NilsandClaude Fable 5 fb5fb13806 tied-alpha outcome: all predictions confirmed; free B causally isolated as fidelity culprit
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 04:13:11 +02:00
NilsandClaude Fable 5 0a5cd80dc1 halting probe: output-stable orbit, not state fixed point; no free ACT at state level; paper claims softened
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 03:29:26 +02:00
NilsandClaude Fable 5 8981cda31d phase diagram: add per-depth point (fidelity kept, gain depth-stranded)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 03:25:16 +02:00
NilsandClaude Fable 5 b53538e82c per-depth outcome scored (depth-stranded content, shared-tail rescue); kcurves updated with k=8
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 03:20:49 +02:00
NilsandClaude Fable 5 82cc1b85b8 depth-curve small multiples per regime (fig_kcurves)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 03:16:58 +02:00
NilsandClaude Fable 5 0e6d6adcb9 fidelity factorial arms (randk / noises0 / hidden-capacity) + pre-registration item 15
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 02:47:14 +02:00
NilsandClaude Fable 5 d39e02f4b9 parcae outcome scored: dynamics confirmed, fidelity refuted -- stability and fidelity are independent dials
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 02:38:51 +02:00
NilsandClaude Fable 5 16f4a057aa TiedAlphaAdapter: learned per-dim alpha with tied B (anchored by construction); pre-registration item 14
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 02:33:14 +02:00
NilsandClaude Fable 5 f370883e23 phase diagram: stability vs fidelity are independent dials (fig_phase)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 02:24:56 +02:00
NilsandClaude Fable 5 1076d57d96 rec arm outcome: rho->4.5 (norm-projected churn), prediction (a) confirmed through k=8; k16/32 cancelled
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:39:22 +02:00
NilsandClaude Fable 5 3477fb5da0 method: truncated-BPTT bias bound via rho(A) — contraction certifies tail-only gradients
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:08:58 +02:00
NilsandClaude Fable 5 d00f120f27 PerDepthAdapter (Bae-style per-iteration merges) + convergence-halting probe (free ACT); pre-registration item 13
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:07:01 +02:00
NilsandClaude Fable 5 f254b20b62 ParcaeAdapter: rho(A)<1 by construction (ZOH negative-diag), rho logging in rec arms, pre-registration item 12
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 00:53:34 +02:00
NilsandClaude Fable 5 d2da8044c3 RecurrentAdapter arm: Huginn-regime retrofit (learned A/B, noise h0, randomized depth) + pre-registration item 11
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 00:49:44 +02:00
NilsandClaude Fable 5 0fd93cb328 n=500 controls: paired net-effect significant (loop vs untrained merge 17-4, p=0.007); ladder and net accounting finalized
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 20:08:57 +02:00
NilsandClaude Fable 5 1d4cffd6ad integrate LCB dissociation (untrained merge wins far-transfer, p=0.001 vs trained loop), capability panel (no MC damage), L23 exit completes sweep
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:44:07 +02:00
NilsandClaude Fable 5 caef95f10f budget-CoT per-item: paired ties with latent arms confirmed (p>=0.69)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 17:29:56 +02:00
NilsandClaude Fable 5 2ce646d175 best-of-3 oracle/deployable split: realistic selector loses the overall edge (55.0/27.3 vs oracle 57.8/34.5); paper economics rewritten
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 17:16:00 +02:00
NilsandClaude Fable 5 e5fc453005 pin gemma-4 model revisions (upstream 12B template regression: thought channel -> empty generations)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 17:13:35 +02:00
NilsandClaude Fable 5 378bb36b8c node_setup: retry+timeout around model downloads (xet mid-transfer stalls)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 16:56:17 +02:00
NilsandClaude Fable 5 7821f304f8 LiveCodeBench transfer harness: raw-jsonl loader (script datasets dead), stdin judge, newest-window selection, plan labeling
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 16:48:12 +02:00
NilsandClaude Fable 5 dcd0b0b2dc HumanEval transfer: trained-vs-untrained-merge paired test (p=0.80, n.s.) -- transfer gain is the merge itself, not trained content
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 16:30:36 +02:00
NilsandClaude Fable 5 d1bafc2d4d revision per review: seed means in headlines, net accounting vs untrained merge, bucket definition up front, GSM12B n.s. (p=0.86), law->regularity, norm spec, compute appendix, public repo URL
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 16:05:40 +02:00
NilsandClaude Fable 5 01fd272028 panel: chat-template MC scoring (raw scoring is chance on -it model), parquet dataset mirrors, incremental dump
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 16:00:17 +02:00
NilsandClaude Fable 5 87308cfb0b score outcomes against pre-registered protocol items 1-10
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 15:48:36 +02:00
NilsandClaude Fable 5 55665ce251 CPU endgame: stats pass (Wilson/McNemar/pooled hard), final figures, PAPER.md rewrite around amortizable-content thesis
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 15:46:58 +02:00
NilsandClaude Fable 5 6a954bd3e4 MC capability panel: ARC/WinoGrande/HellaSwag/MMLU with prompt-loop
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 13:58:57 +02:00
NilsandClaude Fable 5 b461970c9a warm-start flag (stacking arm), adaptive flags for Blocksworld
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 13:02:46 +02:00
NilsandClaude Fable 5 aa3e11fb20 gpuq worker: tee job output to pane for live viewing
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 13:00:38 +02:00
NilsandClaude Fable 5 f2cbce5a03 related_work: log 12b_adaptive outcome as confirmation of the anchoring account
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:57:08 +02:00
NilsandClaude Fable 5 e152a91c05 parse_known_args: module-level parser must tolerate importers' argv
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:50:45 +02:00
NilsandClaude Fable 5 cd8e94dda5 gpuq: idle-time class dispatcher (single-writer, race-free) + docs
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:47:23 +02:00
NilsandClaude Fable 5 9329f22a63 gpuq: class-based scheduling at submit time (pending count + live util)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:41:34 +02:00
NilsandClaude Fable 5 c718c636a5 GPUQ: as-deployed cheat sheet + job conventions
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:40:39 +02:00
NilsandClaude Fable 5 d2d372c909 feedforward flag for HumanEval eval
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:34:57 +02:00
NilsandClaude Fable 5 12ec01becc adaptive flag for GSM eval
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:33:03 +02:00
NilsandClaude Fable 5 95be35d20f adaptive flag for unified trainer
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:32:36 +02:00
NilsandClaude Fable 5 32728e7b07 bootstrap via public https clone (repo now public)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:25:07 +02:00
NilsandClaude Fable 5 471cf336d0 gpuq: document the scheduler (design, contract, ops, failure modes, alternatives)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 11:54:38 +02:00
NilsandClaude Fable 5 0252641701 gpuq: per-node supervisor (worker resurrection + bucket heartbeat)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 11:50:10 +02:00
NilsandClaude Fable 5 5efe5b790a gpuq: bucket-backed per-GPU job queue (worker loop + submit/status helper)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 11:48:03 +02:00
NilsandClaude Fable 5 1460cf241f GSM plan-distillation control (width/depth law falsification test)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 11:22:46 +02:00
174 changed files with 129049 additions and 284 deletions
+98
View File
@@ -0,0 +1,98 @@
# Session handoff — 2026-07-16 ~09:45
Previous session ended because the auto-mode permission classifier entered
a persistent session-level lockout (blocked nearly all Bash/Write on
"earlier conversation content" grounds, ~7h intermittent). Everything
below is current as of ~09:45. Repo knowledge map: `INDEX.md`. Scored
experiment record: `results-loop/PROTOCOL_UNIFIED.md` (items 121, all
scored except pending stats below). Research plan: `PLAN_SELFPACED.md`.
## RUNNING RIGHT NOW — do not disturb, monitor these
### Rented H100 node ("node 1")
- **ssh -p 52525 root@185.151.171.35** (vast.ai; greets with a banner —
filter it). Costs money while it runs.
- **`/root/pipeline5.log`**: lens-mapping pipeline (setsid, survives ssh
drops). Sequence: 31B jbar DONE (saved+synced) → 31B exp4 scan DONE as
a printed TABLE (its `regimes.pt` save crashed — known bug, see
fix-list) → 31B weights cleaned → **26B-A4B jacobian scan RUNNING**
(started ~07:40, ~40-90s/prompt × 256; expect done late morning).
After it: 26B exp4 scan (same save bug will strike — harmless, the
table survives in the log), CLEANED, PIPELINE-DONE.
- **`/root/node_sync.sh` sidecar** (also setsid): rclone-pushes
`/workspace/jspace/results/` (jbar_*.pt, exp4_*.log) and pipeline5.log
to `jspace:jspace/results-lens/` every 120s. The node HAS rclone
credentials (Nils placed them himself — never copy credentials to
nodes from automation; hard-blocked + it's his call).
- When PIPELINE-DONE: verify `results-lens/` has jbar_31b.pt,
jbar_26b_a4b.pt, exp4_31b.log, exp4_26b_a4b.log — then the node can be
destroyed (tell Nils; it's his dashboard).
- Storage lore: /workspace = disk (survives stop, dies on destroy);
/dev/shm = RAM (dies on stop). See memory/vast-node-storage.md.
### Spark (this machine)
- gpuq worker in tmux session `gpuq_spark:gpuq0`, queue EMPTY, healthy.
Worker code now has: data contract (`# gpuq-in:`/`# gpuq-out:` job
comments → auto-sync inputs/outputs with the bucket, outputs every
120s during the job) and graceful drain (`touch ~/gpuq/spark-gpu0/STOP`
— NEVER kill the tmux session mid-job; that cost 3.5h once).
- Thermal: box hard-froze 4× on 2026-07-15 under sustained load —
cabinet now open, stable since. If it freezes: reboot, restart worker
(`tmux new-session -d -s gpuq_spark -n gpuq0 'bash
~/jspace/scripts/gpuq_worker.sh spark-gpu0 0'`). NO @reboot cron
(Nils vetoed — crash-loop risk).
## IMMEDIATE PENDING (blocked by the classifier, run first)
1. `.venv/bin/python scripts/mcnemar_carrycot.py` — the paired p-values
for last night's headline result, especially the drop-bucket test.
THE morning number; item 21's outcome note says "pending".
2. Commit everything uncommitted:
`git add -A scripts/ PLAN_SELFPACED.md HANDOFF.md && git add -f
results-loop/PROTOCOL_UNIFIED.md results-loop/eval_gsm_carrycot*.json
results-loop/gsm_cot_data.json && git commit && git push origin main`
(uncommitted: mcnemar script, lens-noise trainer changes
[train_carry_cot.py --lensnoise, UNVERIFIED — parse-check first],
plan E2-L/E2-A2/E2-N sections, protocol item-21 scoring, this file).
3. Parse the 31B (and later 26B) regime tables from exp4_*.log into
regimes_*.pt — copy the pattern in `scripts/parse_e4b_regimes.py`
(adjust layer count: 31B=60, 26B=48? read from log). Then fill
REGIMES.json rows (results/REGIMES.json) and re-render
`scripts/fig_regimes.py` (5-scale figure, panels auto-fill).
4. Patch `scripts/exp4_regimes.py` to take an explicit output path
(torch.save to CWD-relative "results/" has now crashed 3 scans).
## LAST NIGHT'S RESULTS (already scored in the protocol)
- **Item 21, the headline**: GSM8K carry-cot (dense self-distilled terse
scratchpads through the carry whiteboard): **57.4%** best cell vs
12.1% all-time prior best; base 10.9. Control (same supervision, no
recurrence): 54.7 best. Whiteboard's specific edge: DROP items
(+9/+14 across cells) — reach into problems unreachable at labeling.
If McNemar confirms → gates open for: stage B internalization ladder
(PLAN E2-L), A2 on-policy refresh (E2-A2), lens-shaped noise (E2-N,
Nils's idea, N1 already implemented as --lensnoise, unverified).
- **Item 20**: gate threshold curve — E0's frozen logistic probe
dominates every learned gate; ORACLE gate = 59.6% overall at 0.24
mean iters (hard items depth-diverse: 18/28 solvable at some k, ≤13
at any single k). Gate program continues; binding constraint =
classifier quality on the pre-loop state.
- **E4B lens surprise** (REGIMES.json, provisional): no E2B-style
workspace signature — sensor L11-22, motor from L23, persistence bump
ABSENT; candidate band sits inside the KV-shared zone. The elastic
pair (E2B⊂E4B) reorganized. Needs ignition cross-check + jbar top-up
before strong claims.
## MORNING DECISION QUEUE (Nils decides, one submit each)
With McNemar in hand: which of stage B (internalization) / A2
(re-harvest) / E2-N1 (lens-noise, verify parse first) gets the Spark.
All pre-registered or planned in PLAN_SELFPACED.md; job template
pattern: scripts/jobs/zzz_l_gsm_e2a.sh (uses the data contract).
## OPERATIONAL CAUTIONS (paid for in blood, see LESSONS.md 1-12 + memory)
- pkill -f self-match; k=0 sanity row in every eval; e400 pre-commit;
seeds before believing single cells (noise-s0 taught this twice);
monitors: grep patterns must match eval_loop.py's "acc=" (GSM) vs
"pass@1=" (MBPP); artifacts leave nodes within one sync cycle.
+84
View File
@@ -0,0 +1,84 @@
# Where to find what
Two projects share this repo: the **J-lens reproduction** (does the 2026
workspace paper replicate on gemma-4-E2B?) and the **workspace-looping
investigation** that grew out of it (retrofit recurrence onto the lens-found
band; what does it actually buy?). The second is the active one.
## The claims and their evidence
| you want | look in |
|---|---|
| Current claims, all numbers, figures | `PAPER.md` (source of truth; the .pdf snapshots lag it) |
| What worked / what failed / design rules / ops pitfalls | `LESSONS.md` — read before running anything on this hardware |
| Pre-registrations + scored outcomes (17 items, incl. refutations) | `results-loop/PROTOCOL_UNIFIED.md` — the methods backbone; every claim in PAPER.md §3 traces to an item here |
| Significance tests behind any claimed number | `results-loop/STATS.md` |
| The v2 prototype plan (self-paced workspace: learned gating) | `PLAN_SELFPACED.md` |
| Lab-notebook narrative of the looping investigation | `WORKSPACE_LOOPING.md` (superseded where it disagrees with PAPER.md) |
| Base-reproduction results (lens replication itself) | `RESULTS.md`, `README.md` |
| Per-model lens maps: workspace bands, KV-share boundaries, pinned revisions | `results/REGIMES.json` (canonical registry) + `results/jbar*.pt` (raw J̄) + `results/exp4*.log` (regime scans) |
## Code (`scripts/`, `jlens/`)
- `jlens/core.py` — the lens: model loading (`JLENS_MODEL` env), J̄ readouts.
- `scripts/loop_common.py` — everything band-looping: `BandLooper`
(capture/re-run machinery, KV-cache-safe), `generate_frozen_prompt`
(the ≥3.5× deploy path), and every adapter variant from the regime sweep
(`MergeAdapter` ★, `AdaptiveMergeAdapter`, `RecurrentAdapter`,
`ParcaeAdapter`, `NoisyMergeAdapter`, `TiedAlphaAdapter`,
`PerDepthAdapter`). Band via `JLENS_BAND` env (default E2B 14,30).
- Trainers: `train_merge_code.py` (MBPP; all regime flags live here),
`train_merge.py` (GSM, old full-position regime — historic),
`train_merge_unified.py` (multi-task, hardened protocol),
`train_merge_bw.py` (Blocksworld), `train_distill*.py` (plan distillation).
- Evals: `eval_loop_code.py` (MBPP pass@1 vs k; per-item logs; `--halt`),
`eval_loop.py` (GSM), `eval_bw.py`, `eval_humaneval.py`, `eval_rust.py`,
`eval_lcb.py`, `eval_mc_panel.py`, plus `prep_*.py` (STaR labeling).
- Figures: `fig_*.py` regenerate the canonical PNGs from the JSONs.
- Infra: `gpuq_*.sh` + `GPUQ.md` (bucket-backed GPU job queue),
`node_setup.sh` (vast.ai bootstrap; pins model revisions — see LESSONS #12),
`vast-ai-notes.md`.
## Results directories — including the honest mess
- `results-loop/`**the looping project's data**: 84 `eval_*.json`
(tag suffixes: `_s<seed>`, `_rec16`/`_parcae16` recurrent arms, `_pd4`
per-depth, `_ta` tied-alpha, `_rk16` random-depth, `_ns` noise-s₀,
`_h2048` capacity, `code2gsm_*` transfer; `per_item` only in files from
Jul 14 onward), adapter checkpoints (`adapter_*.pt`, e400 = the
pre-committed eval checkpoint), canonical figures (`fig_kcurves.png`
design-space grid, `fig_phase.png` two-dials diagram, `fig_loop_vs_ff.png`
recurrence-vs-distill ladder, `fig_placement/transfer/scale.png`),
and `chain*.log` — autonomous-session logs, archaeology only.
- `results/` — lens reproduction outputs + the cross-model registry
(`REGIMES.json`).
- `results-band-*/` — one directory per entrance-placement arm of the
placement sweep (L2L24 entrances); summarized in PAPER fig_placement;
kept for per-item audit.
- `results-12b/`, `results-loop-12b/` — 12B lens map and looping arms.
- `results-tap23/`, `results-tap34/`, `results-kvtest/`, `results-combo/`,
`results-panel/`, `results-distill-s7/` — single-question side arms
(exit-tap sweep, KV nulling check, combined arms, MC panel, distill seed).
- `results-node*/`, `results-node2-final/` — raw syncs from rented H100
nodes (500-item eval campaign).
- `results-26b/`**unclear provenance** (Jul 13; layer indices ≤26 mean
it is NOT the 26B MoE despite the name — possibly a misnamed early scan).
Trust nothing here without re-derivation.
- `paper-A/`, `paper-B/`, `paper-D/` — abandoned paper-outline variants
(one PLAN.md each); the live outline is PAPER.md itself.
- `related_work/` — the two anchor papers (McLeish 2511.07384,
Lys 2602.14759), the workspace paper, `relevant_to_us.md` notes,
`bibliography.bib`.
## Conventions worth knowing
- Every eval prints a `k=0` row first; it must equal the base model
bit-exactly (0.488 on MBPP-250) — the sanity anchor that has caught two
silent bugs (LESSONS #2, #12).
- Difficulty labels (`easy`/`hard`/`drop`) are STaR self-labels:
direct-pass / CoT-only-pass / unreachable. "hard" = plan-dependent.
- Checkpoints are pre-committed before evals (usually e400); post-hoc
checkpoint shopping is flagged as exploratory wherever it happened.
- GPU jobs go through the gpuq queue (`gpuq_submit.sh <worker> <job.sh>`),
never bare nohup on the Spark; jobs are killed by `pkill -f` self-matches
embarrassingly often (LESSONS #6).
Binary file not shown.
+19 -1
View File
@@ -62,7 +62,15 @@ reproduction: [`RESULTS.md`](RESULTS.md). Everything on `google/gemma-4-E2B-it`
6. **Background jobs must be `setsid`'d** or the harness/session restart
kills them mid-run. And `pkill -f <pattern>` will match your own launcher
shell if the pattern appears in its command line.
7. **Zero-init adapter output layer ⇒ zero grads upstream at step 0** — on
7. **Sustained training in a closed cabinet = thermal hard-freezes.** Four
crashes in one day (journal stops mid-line, no OOM, no shutdown trace,
37GB free at one death) on a DGX Spark that was stable all week under
light load. Pattern: only under hours of continuous GPU load; fixed by
opening the cabinet. Diagnose by exclusion: earlyoom quiet + journal
truncation + load correlation = thermal, not software. And do NOT
auto-restart training via @reboot cron on a thermally-suspect box — it
risks a crash loop with no human circuit breaker.
8. **Zero-init adapter output layer ⇒ zero grads upstream at step 0** — on
`mlp[0]` this is expected (LoRA-B-style), not a bug; check the output
layer's grad instead.
@@ -76,3 +84,13 @@ reproduction: [`RESULTS.md`](RESULTS.md). Everything on `google/gemma-4-E2B-it`
3. **Verify with the lens, gate with the labels**: the J-lens picks the band,
measures whether loops compute, and diagnoses failures; STaR difficulty
labels supervise both the curriculum and (next) the adaptive-depth gate.
## 12. Pin model revisions on fresh nodes
Upstream updated google/gemma-4-12B-it mid-project: the new chat template
adds a `<|channel>thought` scaffold, and greedy generation closes the empty
thought channel and stops — every generation decodes to "". Symptom:
0/500 pass rates with rc=0 (looks like a harness bug, is a silent model
swap). E2B was unaffected. Fix: `hf download --revision <hash>` + repoint
`refs/main` in the cache; node_setup.sh now pins both models (12B
0e2b1058…, E2B 9dbdf8a8…). Rule: any cross-node result assumes identical
model revisions — pin them, don't trust "main".
+468 -235
View File
@@ -1,286 +1,519 @@
# Retrofitting Latent Planning onto a Frozen Language Model via Workspace Recurrence
# Latent Planning by Workspace Recurrence: an Interpretability-Placed Implant, and What It Actually Buys
*Working draft, 2026-07-14. All experiments: google/gemma-4-E2B-it (frozen), single DGX Spark. Code and artifacts: `~/jspace`.*
*Revision draft, 2026-07-14. Base models: google/gemma-4-E2B-it and
gemma-4-12B-it, both frozen. Hardware: DGX Spark + rented 2×/8×H100 nodes.
Code, per-item logs, and pre-registrations:
https://git.draic.info/nils/jspace (public). Statistics:
`results-loop/STATS.md`.*
## Abstract
Interpretability work with an averaged-Jacobian lens ("J-lens") shows that
mid-depth layers of a pretrained language model form a *workspace*: a band of
layers that holds verbalizable, unspoken intermediate content. We ask whether
that band can be **iterated in place** — spending more serial compute per
input without emitting reasoning tokens — on a *frozen* model. A naive loop
diverges: the band is not a self-map. We show that a 1.6M-parameter
**anchor-dominant merge adapter** (0.03% of the model) at the band entrance
makes the recurrence a stable fixed-point iteration, and that training only
this adapter — with self-generated, verifier-filtered supervision and a
difficulty→depth curriculum — turns iteration into computation. On MBPP,
looping the workspace over the prompt ("latent planning") raises pass@1 on
plan-dependent problems from **5.5% to 30.943.6%** (three seeds, full test
set, execution-verified); overall accuracy is unchanged-to-slightly-improved
(51.8% → 51.853.8%, within noise at n=500) — the method's value is
cost-shaped (silent, prefill-parallel, no per-token overhead), not
accuracy-dominance. Controls attribute the hard-bucket gain to the
recurrence itself: a same-size adapter trained on identical data *without*
the loop reaches only 17.9%, exactly matching the untrained loop. On GSM8K the picture inverts — no recurrent variant beats the
weights-only control — and a four-arm decomposition localizes why: the loop
performs *plan refinement*, which code synthesis needs and answer-time
arithmetic does not. The J-lens provides both the intervention's design
(where to loop) and its verification (latent concepts sharpen ~8× per
converged iteration). Because the looped prompt states are constant during
generation, latent planning is prefill-shaped and adds no per-token cost.
Interpretability work with an averaged-Jacobian lens ("J-lens") partitions a
pretrained language model's depth into regimes, including a mid-depth
*workspace* band that holds verbalizable, unspoken intermediate content. We
retrofit recurrence onto this band in a **frozen** gemma-4-E2B: a
1.6M-parameter anchor-dominant merge adapter (0.03% of parameters) at the
band entrance turns the non-self-map band into a stable recurrence, trained
with self-generated, verifier-filtered supervision. Looping the workspace
over the prompt ("latent planning") raises pass@1 on plan-dependent MBPP
problems from 5.5% to 37.5±5.5% over five seeds — pooled across MBPP,
HumanEval, and Rust/MultiPL-E, 4.2%→35.6% (McNemar p≈1.5e-10) — with zero
visible tokens and zero decode cost. Placement is decisive, not convenient:
the gain exists only at the lens-identified boundary (L14), collapsing below
it, and KV-sharing structurally nulls entrances above it. Net of the
untrained-merge perturbation floor (20.0%), the loop-specific effect
survives paired testing (p=0.007).
## 1. Introduction
A two-part attribution program then bounds the mechanism. First, the
content is *amortized, not computed*: recurrence-free plan-distillation
into the same adapter matches the loop, gains do not stack, and inference
depth beyond k≈4 is flat — the state trajectory is an output-stable orbit,
not a converging computation (half of prompts' states never converge at
cos 0.9995 by k=8, with no difficulty gradient, so no convergence-based
early exit falls out). Width rivals depth on code (trained pause registers:
36.4%); recurrence is needed where state must evolve (GSM8K carry,
Blocksworld planning). Second, a pre-registered regime sweep spanning
unconstrained learned recurrence (Huginn-style), spectrally constrained
state maps (Parcae-style), per-iteration weights (Bae-style), and learned
anchor coefficients shows that **dynamical stability and substrate fidelity
are independent dials**: spectral radius governs convergence only (an
unconstrained map drifts to ρ≈4.5 with no fit benefit; a constrained one
stays at ρ≈0.3 with no fit cost — both lose 17 points of easy-item
accuracy), while fidelity is governed by fixed-point *location*, causally
isolated to one design choice — tying the input map to the anchor's convex
complement, B=(1−α)I. Per-iteration weights strand the gain at trained
depths; every regime buys the same hard-bucket gain (3646%); no regime
exceeds the amortization ceiling at this budget. The hand-tuned recipe is
thus the measured optimum of its design space, not a lucky point in it.
At 12B the anchor coefficient must become state-dependent (3.8K parameters)
to preserve the substrate — the one dial that is task- and scale-dependent.
Details and exact numbers: §1 and §3.
Large language models buy reasoning accuracy with emitted tokens: chains of
thought give the network more serial passes, at the cost of latency, output
tokens, and bandwidth-bound decode. Recurrent-depth architectures (Universal
Transformers; DEQs; Huginn, arXiv:2502.05171; Mixture-of-Recursions,
arXiv:2507.10524) buy the same serial compute silently — but require
(pre)training the recurrence in at scale.
We investigate a middle path: **retrofit** recurrence onto an off-the-shelf
frozen model, using an interpretability signal to decide *where*. The
J-lens (from the "verbalizable global workspace" line of work) partitions
depth into transduction, sensor, workspace, and motor regimes; the workspace
band (L1430 of 35 in our subject model) holds slowly-varying, unspoken
intermediates — e.g. 'spider' before answering "8" to *"the animal that spins
webs has how many legs?"*. If the workspace approximates "iterate toward a
settled representation", looping it should deepen computation without
parameters. The contributions:
## 1. What this paper claims
1. **A minimal retrofit that works**: an anchor-dominant merge
(`(1−α)e + α·ŝ + MLP([e;ŝ])`, α=0.3, MLP zero-init, 1.6M params) makes
the frozen band a stable, answer-preserving recurrence; training only the
merge makes iterations *sharpen* rather than hold.
2. **A verified capability gain** on plan-dependent code synthesis, with the
full attribution grid (weights / untrained loop / trained loop / pause
tokens) showing the recurrence is the active ingredient.
3. **A mechanistic boundary**: math inverts the result, and the decomposition
(prompt-side vs generation-side × weights vs recurrence) identifies the
mechanism as plan refinement, not generic extra compute.
4. **Deployment properties**: bit-exact KV-cache-compatible inference (loop
once at prefill), a difficulty gate trained free from the labeling
pipeline, and economics that improve with model scale.
*(One model family, two scales: we state findings as empirical regularities,
not laws.)*
1. **A placement regularity.** The retrofit works if and only if the
recurrence enters at the lens boundary. Entrances at L9L13 (same
adapter, data, curriculum) destroy overall accuracy (1434% vs 52%)
while recovering at most half the hard-bucket gain; entrance at L14
preserves overall and maximizes the gain (fig_placement). Entrances at
L17/L24 are *structurally null* in this architecture: KV-sharing makes
layers ≥15 reuse keys/values computed at ≤14, so k>0 is bit-identical to
k=0 — a hazard for any retrofit method that skips the mechanistic check.
Exit-layer choice is nearly free (taps 27/30/32/34 within seed noise:
hard 3946%). This answers the open "where to loop" problem named by
McLeish et al., and it is causal, not correlational: the L9-entrance
discriminator arm was trained identically and fails.
2. **A verified, statistically solid capability gain on a narrow slice.**
Plan-dependent items (the model solves them with an explicit written plan
but not directly): seed-mean 37.5±5.5 on MBPP (best 43.6%); pooled across
three benchmarks, 4.2%→35.6%, p≈1.5e-10. Overall accuracy is
statistically unchanged on MBPP (p=0.34) and improved on HumanEval
transfer (58.5%→66.5%, p=0.011). Net of the untrained-merge floor
(20.0% at n=55), the loop-specific effect is +17.5 points (seed mean)
and **survives the paired test** (loop vs untrained merge on hard,
p=0.007; distill vs untrained, p=0.0075).
3. **A deflationary mechanism finding.** The trained loop converges to a
fixed point by k≈34 and behaves as *amortized plan content*, not
iterative computation: plan-distillation into the identical architecture
without recurrence matches it; stacking buys nothing (loop-training a
distill-warmed adapter: 34.5%, below distill alone; running the distilled
adapter in loop mode: drops to 20.0%); deeper k at inference is flat
(k=8: 40.0%; output-stable despite residual state drift, §3.8). The
recurrence is a *training-time scaffold* that lets the
adapter find plan-shaped content — content that can equally be put there
by distillation if plans are available.
4. **A width-vs-depth pattern.** Trained pause registers (width) capture most of
the plan effect on code; recurrence (depth) is needed only where a state
must *evolve* — on GSM8K generation-side carry beats registers, and on
Blocksworld (pure planning, no world knowledge) the loop lifts hard-split
plans 0%→43% at 2B where everything else fails. Plans are wide; execution
is deep.
5. **Honest economics.** The implant's costs: ≈2.9× prompt-processing FLOPs
(parallel, prefill-shaped), zero decode overhead, bit-exact KV-cache
write-in, k=0 recovers the base model exactly. Its competition at
comparable compute (accounting in Appendix A — FLOPs, wall-clock, and
token budget do not rank the arms the same way): *oracle* best-of-3
(verifier-assisted) wins overall accuracy (57.8%, paired p=0.020 vs
loop), but the *deployable* logprob-selected variant drops to 55.0%
overall / 27.3% hard — indistinguishable from the latent arms overall
and directionally behind on hard. A 50-token visible plan ties the hard
bucket. The value proposition without a verifier: no visible tokens, no
decode latency, and the hard-slice specialization (distill's 46% >
budget-CoT's 40% > deployable sampling's 27%).
6. **Scale transfers only with a state-dependent stability dial.** At 12B the
2B-tuned constant α=0.3 collapses overall accuracy (72.6%→43.0%); the
damage is present *before* adapter training (untrained-loop arm) and is
not fixed by retuning α or LR. A per-position learned coefficient
α=σ(w·[e;ŝ]+b) restores MBPP (overall 69.4%, hard 11.4%→27.3%) — but
fails to rescue Blocksworld-12B and yields only a nominally positive,
not significant GSM8K-12B overall delta (35.9%→36.7% at k=1, p=0.86).
## 2. Method
**Locating the band.** The lens reads residual state h at layer through the
averaged Jacobian J̄_ = E[∂h_final/∂h_] and the unembedding. Depth regimes
follow from what the readout tracks (input echo / abstract content / output
token). On gemma-4-E2B: workspace ≈ L1430 (1.08B params, 58% of decoder).
averaged Jacobian J̄_ = E[∂h_final/∂h_] and the unembedding; depth regimes
follow from what the readout tracks. On gemma-4-E2B: workspace ≈ L1430 of
35; on 12B: L3645 of 48.
**Making the band a self-map.** Feeding L30's output to L14 collapses in one
step (out-space ≠ in-space; norms and content 17 layers "downstream").
Additive anchoring diverges. The fix is DEQ-style input injection done by
hand: with e = L13's output (fixed anchor) and s the fed-back band output,
**Making the band a self-map.** Feeding L30's output back to L14 collapses
(out-space ≠ in-space). With e = L13's output (fixed anchor) and s the
fed-back, norm-matched band output:
L14-in = (1−α)·e + α·(s · |e|/|s|) + MLP([e ; s·|e|/|s|]), α = 0.3.
L14-in = (1−α)·e + α·ŝ + MLP([e ; ŝ]), ŝ = s·‖e‖₂/‖s‖₂
Zero-initializing the MLP's output layer makes the untrained adapter exactly
the hand merge, which is stable and answer-preserving for ≥11 iterations but
only *holds* content (lens concept flat).
(per-position L2 norms over the hidden dimension, computed in fp32)
**Training only the merge.** Supervision is self-generated and
verifier-filtered (STaR-style): the frozen model attempts each task directly
and with explicit planning/CoT; items it solves only with planning are
"hard", direct solves "easy", neither "drop". Cross-entropy on answer/code
tokens of the *direct* prompt; the model's own verified outputs are the
targets (in-distribution). A **difficulty→depth curriculum** trains easy
items at loop depth k=1, mixed at k=2, hard only at k=24, so loss on hard
items is reducible only through the recurrence. For generation tasks the
loop applies to the **prompt span only** ("latent planning"): generated
tokens run the plain path but attend to the looped prompt states; this
removes exposure bias structurally.
α=0.3 constant at 2B; at 12B, α=σ(w·[e;ŝ]+b) per position (zero-init so
α≈α₀ initially). MLP output zero-init: the untrained adapter is exactly the
hand merge — stable, answer-preserving, content-holding.
**Inference cost.** Causality makes the looped prompt states independent of
generated tokens, so they are computed once; a hooked prefill writes them
into the KV cache and generation proceeds natively (verified bit-identical;
≥3.5× faster than recomputation). The concrete overhead at k=4 is 5 passes
over the band's 17/35 layers at prefill — ≈2.9× prompt-processing FLOPs,
parallel across positions — and **zero** additional decode cost. Explicit
planning with ~200 emitted tokens costs more total FLOPs and pays them
serially at bandwidth-bound decode; this asymmetry grows with model size.
**Training.** STaR-style self-labeling: items the frozen model solves only
with an explicit plan/CoT are "hard", direct solves "easy", neither "drop".
Cross-entropy on answer/code tokens of the direct prompt, model's own
verified outputs as targets. Difficulty→depth curriculum (easy k=1, mixed
k=2, hard k=24). The loop applies to the **prompt span only**; generated
tokens run the plain path but attend to looped prompt states. Variants
trained the same way: **pause-N** (N trained register tokens appended to the
prompt, no recurrence), **plan-distill** (KL from the model's own
plan-in-context distribution into the FF adapter), **rung-2** (warm-started
adapter + entrance-faded LoRA rank 8 on the band's first layers, loop-only
via a global toggle), and **stack** arms (distill-warm + loop training;
distilled adapter evaluated in loop mode).
**Truncated backprop is certified by contraction.** Recurrent-regime arms
train with tail-only BPTT (gradients through the last 4 iterations; the
no-grad prefix stores no activations, so memory is constant in depth).
The truncation bias scales as ρ(A)^(Ttail) — at ρ=0.3 the discarded terms
are ≤1%, making the cheap estimator essentially exact; at ρ≥1 it is
dominated by what it discards. Stability, fixed-point convergence, valid
tail gradients, and the convergence-halting exit signal are all the same
dial.
**Inference.** Looped prompt states are causally independent of generated
tokens: computed once at prefill, written into the KV cache by a hooked
forward pass, generation native. Verified bit-identical to the slow path.
Cost at k=4: ≈2.9× prefill FLOPs, **zero** decode overhead.
## 3. Results
### 3.1 Latent planning on code (MBPP)
Statistics throughout: Wilson 95% CIs; paired comparisons by exact McNemar;
all headline arms evaluated on the full 500-item MBPP test split (hard
bucket n=55), HumanEval n=164 (hard n=38), Rust/MultiPL-E n=154 (hard n=25),
execution-verified.
Full 500-item test split, greedy decode, unit-test-verified. Hard bucket =
items the frozen model solves only with an explicit written plan (n=55).
**Bucket definition, stated up front.** Hard labels come from labeling runs
of the frozen base model on the test items themselves (direct vs
plan-in-context, greedy). This is legitimate for *descriptive* slicing but
would be circular for selection — so no arm, hyperparameter, checkpoint, or
loop depth was ever chosen using bucket results (pre-registered;
`PROTOCOL_UNIFIED.md` items 12, 8). Because the bucket conditions on k=0
failure, regression-to-mean inflates *any* intervention's bucket score: the
untrained merge already reaches ~18%, and we therefore report the
loop-specific effect **net of that floor** wherever attribution is claimed.
Robustness: redefining "hard" as labeled-hard ∧ k=0-fails-in-all-five-seeds
(52/55 items) moves headline numbers <2 points; both definitions share the
base model, which an independent difficulty proxy would not — we flag this
as an open external check.
| k=4 (prompt-only loops) | hard pass@1 | overall |
### 3.1 The placement regularity
![Placement cliff](results-loop/fig_placement.png)
Entrance-layer sweep with everything else fixed. L14 (lens boundary):
hard 43.6%, overall 53.6%. L13: hard 17.9%, overall 34.4%. L9L12: overall
14.030.8% (substrate destroyed). L17/L24 entrances: k>0 ≡ k=0 (KV sharing;
verified bit-identical) — the 12B model has no shared-KV layers, making it
the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34):
hard 39.346.4%, within seed spread. The lens boundary is necessary; the
exit is a free parameter — the completed five-point exit sweep
(L23/27/30/32/34) spans hard 39.346.4% with L23 at the top (46.4% at k=2),
all within seed spread.
### 3.2 The attribution ladder
![Attribution ladder](results-loop/fig_ladder.png)
MBPP hard bucket (plan-dependent, n=55 unless noted):
| arm | hard pass@1 | overall |
|---|---|---|
| baseline (k=0) | 5.5% | 51.8% |
| trained loop, seed 0 | **43.6%** | 53.6% |
| trained loop, seed 1 | **41.8%** | 53.8% |
| trained loop, seed 2 | **30.9%** | 51.8% |
| base (k=0, bit-exact) | 5.5% | 51.8% |
| untrained loop (α-merge only, k=4) | 20.0% | 50.2% |
| trained FF, no recurrence (k=1) | 27.3% | 53.6% |
| pause-16 registers (width) | 36.4% | 55.2% |
| **trained loop k=4** (seed mean, 5 seeds) | **37.5±5.5** (best 43.6) | 53.6% |
| rung-2: + entrance-faded band LoRA (n=28) | 42.9/46.4 (2 seeds) | 51.2/52.4 |
| **plan-distilled FF** (mean, 8 runs) | **45.7±4.6** (best 49.1) | 55.5% |
| budget-CoT (50 visible tokens) | 40.0% | 53.8% |
| best-of-3 sampling (≈matched FLOPs) | 32.7% | **57.2%** |
| explicit plan in context (ceiling) | 94.5% | 59.0% |
Silent loops recover roughly 40% of what explicit planning achieves, at zero
visible-token cost, with no overall regression (the easy-item perturbation
tax, ~9 points, is offset by hard/drop gains; a gate removes most of it,
§3.4).
Significance structure (McNemar, `STATS.md`): loop vs base on hard,
p=5.7e-6; every latent-arm-vs-latent-arm difference (loop vs distill, distill
vs stack) is **not significant** at n=55; loop vs base *overall* is not
significant on MBPP (p=0.34). The ladder's shape is reliable; its fine
ordering is not.
![MBPP pass@1 vs loop depth](results-loop/loop_eval_code.png)
**Net accounting.** The attribution-critical comparison is trained-loop vs
*untrained merge*, not vs base: gross 5.5→37.5 (seed mean), of which the
untrained perturbation floor is 20.0 points — the loop-specific net is
+17.5 (seed mean) / +23.6 (best seed), and the paired item-level test is
decisive (loop-only 17, untrained-only 4, p=0.007; distill likewise
p=0.0075). The trained-FF control (27.3%) sits between floor and loop,
not significantly above the floor (p=0.48): weights alone buy little
without either recurrence or plan supervision. All controls now n=500 /
hard n=55, same harness.
### 3.2 Attribution: the recurrence is the ingredient
### 3.3 The decisive tests: nothing stacks
250-item subset; same data, same 1.6M parameters, same insertion point:
If the loop performed genuine iterative computation, plan-distilled content
plus recurrence should compound. It does not:
| arm | hard pass@1 |
|---|---|
| baseline | 3.6% |
| untrained loop (α-merge only) | 17.9% |
| trained adapter, **no loop** (weights control) | 17.9% |
| trained **loop** | **42.946.4%** |
- **Distill-warm + loop training**: hard 34.5% — below distill alone.
- **Distilled adapter run in loop mode**: hard 20.0%, overall 45.8% —
looping *degrades* the distilled weights.
- **Pause-16 + distill**: hard 30.9% — no width stacking either.
- **Inference depth beyond convergence**: k=8 hard 40.0% ≈ k=4 (fixed point,
cos(sₖ,sₖ₋₁)=1.000 by k≈34).
The weights control lands exactly on the untrained-loop value: ~18 points is
what perturbation-plus-format-alignment buys. The remaining ~28 points
require iterating the band. Post-hoc depth selection is excluded by
pre-registration (k=2 fixed on validation before test numbers existed;
k-curves reported descriptively).
Reading: the recurrence is a **training-time scaffold**. The curriculum
forces hard-item loss to be reducible only through the loop, and what the
adapter learns to inject is plan-shaped content — the same content
distillation installs directly when explicit plans are available. The loop's
distinctive value is that it finds this content *without* plan supervision
(STaR labels only say which items needed plans, not what the plans were).
**Checkpoint selection.** No checkpoint was chosen using test or generation
results. Seed 0's checkpoint (step 399) was fixed at training time from the
validation-CE overfitting inflection, before any generation eval of that
adapter; seeds 15 use step 400 by pre-commitment made before those seeds
were trained. We separately report that validation CE is a poor proxy for
generation accuracy (a checkpoint selected by val-CE on a sibling arm
underperformed a later one), which is why the fixed-step rule is used
rather than per-seed val selection.
### 3.4 Compute-matched honesty
### 3.3 The boundary: math
At approximately matched FLOPs (Appendix A gives the accounting, separated
into FLOPs, wall-clock, and token budget), the token-space comparison
splits into two very different claims:
On GSM8K, *no* recurrent variant beats the weights-only control. The four-arm
grid (hard bucket) decomposes the failure:
| best-of-3 variant | overall | hard | vs loop (paired) |
|---|---|---|---|
| **oracle** (any-of-3 passes; needs a perfect verifier) | 57.8% | 34.5% | beats loop overall, p=0.020 |
| **deployable** (highest mean logprob of 3) | 55.0% | 27.3% | n.s. overall (p=0.47); loop ahead on hard 167 (p=0.09) |
| GSM8K hard | prompt-side only | touches generation |
|---|---|---|
| feedforward weights | **11.8%** | 4.7% (pause-token control) |
| recurrence | 6.38.7% (prompt loop) | 9.4% (cross-token carry) |
The earlier draft's "sampling wins overall" was the **oracle** number — an
upper bound requiring an external verifier that MBPP's own tests provide
but a deployment does not. With the deployable selector (identical seeded
samples, so the comparison is exact), best-of-3 is statistically
indistinguishable from the latent arms overall, *behind* them
directionally on the hard bucket, and pays ≈3× visible tokens and serial
decode for it. Budget-CoT-50 remains the strongest honest token baseline
(53.8% overall, hard 40.0%; per-item rerun 53.8/38.2) — and the paired
tests confirm it is a *tie* with the latent arms on both axes (p≥0.69 vs
loop and distill), at the price of 50 visible tokens and their serial
decode latency. The implant's advantages at matched compute
without a verifier: zero visible tokens, zero decode overhead, and the
hard-slice crown under distillation (46%). Where a task *does* come with a
cheap verifier, oracle-style sampling is the better overall-accuracy
spend — both halves belong in the deployment picture.
Orthogonal effects: perturbing free-running generation positions is costly
for either mechanism; recurrence beats weights only where a state must
evolve (the generation side — carry doubles the pause control in-harness),
and loses on the static prompt side. No variant beats the 10.5% overall
baseline. Reading: the trained loop performs **plan refinement**; code
synthesis is plan-shaped, multi-step arithmetic is not — its serial
computation happens during the answer, and one frozen band pass per token
cannot perform it silently at 2B. CoT tokens remain load-bearing for math.
(Hard-bucket cells carry an outcome-selection caveat — buckets were defined
by greedy baseline outcomes; sampled relabeling is in progress — so the math
conclusion is stated on overall numbers.)
### 3.5 Width vs depth, and the task boundary
### 3.4 Mechanism and deployment
Pause registers (width) reach 36.4% (16 registers; 8: 30.9%, 32: 34.5% — flat
in N) on MBPP hard: static plan content fits in registers. GSM8K inverts the
prompt-side result entirely (no variant beats the weights control
prompt-side), but generation-side *carry* — recurrence across token steps —
doubles the pause control on hard items: arithmetic's serial state evolves
during the answer. Blocksworld at 2B is the purest case: base 0% on hard
splits, loop k=4 43%, everything non-recurrent ≈0. The pattern: **plans are
wide; execution is deep.** Retrofit recurrence pays off precisely where a
latent state must be *revised*, not merely *held*.
**Fixed point.** The trained loop takes a large first step
(cos(s₁,s₀)=0.926 vs 0.977 untrained) and converges bit-exactly by k≈34
(cos=1.000), where accuracy and lens-sharpening plateau — extra iterations
are no-ops, explaining the k-curve shape.
### 3.6 Scale: the stability dial
![Loop convergence dynamics](results-loop/loop_dynamics.png)
![Cross-scale grid](results-loop/fig_scale.png)
**Lens verification.** P(latent concept) under the J-lens at the band exit
rises 0.015→0.13 across iterations after training (~8× the untrained
control, which only holds). The same lens that located the band verifies
that looping deepens its computation — and makes the silent reasoning
inspectable.
At 12B (no shared KV — unconfounded), constant α=0.3: overall collapses
72.6%→43.0% at k=4 while hard limps to 11.4%. The untrained-loop arm shows
the damage precedes adapter training; α=0.15 and LR retuning do not fix it
(47.6/52.6% overall). The state-dependent coefficient does, on MBPP:
overall 69.4% (base 72.4%), hard 11.4%→27.3%. It does **not** rescue
Blocksworld-12B (easy items destroyed at k=4; constant-α had reached hard
40% but also destroyed easy) and yields a **nominally positive, not
significant** overall delta on GSM8K-12B (35.9→36.7 at k=1; paired McNemar
on 32 discordant items, p=0.86; hard 1.6→10.6) — no arm anywhere in the
program produced a statistically significant overall gain at 12B.
Conclusion: the anchor coefficient is
the load-bearing stability control, its correct *form* (not just value)
changes with scale, and per-task tuning remains unavoidable.
**Gate.** A logistic probe on the k=0 workspace state (supervised for free
by the STaR labels) routes prompts: predicted-easy at k=0, predicted-hard at
k=4. Result: overall equal to the best uniform depth with easy items fully
preserved (97.5% vs 98.4% baseline); probe precision (19% at 64% recall) is
the current ceiling.
### 3.7 Transfer: substrate, not task
**Negative results with content.** Mixed-task (code+math) training regressed
both tasks versus dedicated adapters, despite indistinguishable validation
CE — cross-entropy parity does not predict generation parity. Validation-CE
checkpoint selection likewise failed to track generation accuracy.
![Transfer panel](results-loop/fig_transfer.png)
MBPP-trained implants applied unchanged: **HumanEval** overall 58.5%→66.5%
(loop k=4, p=0.011 vs base; hard 0→31.6%). The decisive control: the
*untrained* merge already reaches 64.6%, and trained-vs-untrained is **not
significant** (paired McNemar at k=2, 9 vs 7 discordant, p=0.80). What
transfers significantly is the *merge perturbation itself*, not the
MBPP-trained content — the cleanest evidence that off-distribution value is
substrate-shaped rather than task-memorized. (The transferred pause adapter
reaches 66.5%, hard 38.9%, consistent with the same reading.)
**LiveCodeBench sharpens this into a dissociation** (150 newest stdin
problems, Nov 2024Apr 2025, execution-verified; no LCB training anywhere
in the pipeline; base 18.7%):
| arm (MBPP-trained where trained) | overall | hard (n=25) | vs base, paired |
|---|---|---|---|
| **untrained merge, k=4** | **24.0%** | **36.0%** | **+**, p=0.039 |
| trained loop, k=4 | 15.3% | 8.0% | , p=0.23 |
| distill FF, k=1 | 12.7% | 16.0% | ****, p=0.049 |
Far from distribution, the *trained content is a liability* (distill
significantly hurts; untrained-vs-trained-loop is 141 discordant,
p=0.001) while the *untrained anchored recurrence significantly helps*
the training-free regime of Lys et al. is the right choice off-distribution,
and the amortized-content reading of §3.3 predicts exactly this: what the
adapter learned is MBPP-shaped plan content, valuable where plans look like
MBPP plans and harmful where they don't. Transfer ordering by distance:
HumanEval (near) — trained ≈ untrained; Rust (mid) — trained helps the hard
bucket; LCB (far) — untrained wins outright. Caveats: single seed per arm,
hard n=25, one benchmark at the far end. **Rust/MultiPL-E** (Python-trained, different
language, compile-run-verified): hard 8.0%→24.0% (p=0.125 at n=25 —
directionally consistent, underpowered). **Blocksworld** MBPP-transfer:
hard 0→14.3% (task-trained: 43%). Content transfers where the substrate's
plan-representation overlaps; task-specific training still dominates.
### 3.8 Mechanism, verification, deployment
The trained loop takes a large first step (cos(s₁,s₀)=0.926 vs 0.977
untrained); accuracy and lens-sharpening plateau by k≈34. A population
probe (n=250, state-cosine threshold 0.9995) shows the plateau is
*output-level*: half the prompts' states are still drifting at 1e-31e-4
cosine scale at k=8 while generation is already depth-stable — an
output-stable orbit rather than a literal state fixed point, with no
difficulty gradient in state-convergence depth. Consequently,
convergence-based early exit ("free ACT") does not fall out of the state
trajectory; halting would need an output-level signal. P(latent concept) under the J-lens at the band exit rises
0.015→0.13 across iterations (~8× the untrained hold) — the lens that placed
the implant also renders its silent content inspectable. The STaR labels
train a free difficulty gate (route predicted-hard to k=4, else k=0);
gate quality (19% precision at 64% recall) is the current ceiling on
removing the easy-item perturbation tax. k=0 is the exact base model by
construction — the implant is removable at token granularity.
**General-capability panel** (ARC-Challenge, WinoGrande, HellaSwag, MMLU;
800 items each, length-normalized MC likelihood via the chat template, loop
applied to the context span). The safety answer is clean — **k>0 does not
damage general abilities**:
| arm | ARC-C | WinoGrande | HellaSwag | MMLU |
|---|---|---|---|---|
| base (k=0) | 36.0 | 55.9 | 52.3 | 30.1 |
| loop k=2 (MBPP adapter) | 36.1 | 55.3 | 49.6 | 31.3 |
| distill FF (MBPP) | 41.8 | 56.6 | 57.0 | 31.8 |
The loop arm is flat within noise (largest move 2.6 on HellaSwag,
unpaired n=800). The distill adapter *nominally improves* every benchmark
(+5.8 ARC, +4.8 HellaSwag) — consistent with §3.7's finding that these
implants carry a generically useful perturbation component, though
MC-likelihood scoring and generation quality are different regimes (see
the LCB result below before reading this as free capability).
### 3.9 Negative results with content
Mixed-task (code+math) training regressed both tasks at equal validation CE
— CE parity does not predict generation parity, and validation-CE checkpoint
selection fails likewise (fixed-step pre-commitment used instead; no
checkpoint was selected on test or generation results). GSM8K distillation
collapsed to empty outputs twice (E2B first attempt, 12B) on 3-token targets
under KL-dominant loss; a CE-dominant retry at E2B trained but reached only
hard 4.7%. Plan-distillation on GSM8K underperforms its MBPP twin even when
training succeeds: consistent with §3.5, there is little static plan content
for math to amortize.
## 4. Related work
Two recent papers bracket this work. **McLeish et al. (arXiv:2511.07384)**
retrofit depth-recurrence into pretrained 1B models via layer surgery +
continued pretraining (~50B tokens, all parameters, Muon, recurrence
curriculum to r=32): the generic claims "retrofitted recurrence works and
beats the non-recurrent parent" and "pretrain-then-convert" are theirs, at
~5 orders of magnitude more training cost than ours. They name layer choice
as an open problem; our lens-derived band with its causal backing (anchor
cliff at L14, tap invariance, wrong-band ≈ 0, KV-sharing hazard) is a direct
answer to it. Unlike their surgery (which needs a healing phase), our k=0
exactly recovers the base model. **Lys et al. (arXiv:2602.14759)** loop
frozen models training-free and show naive looping degrades (distribution
shift) while interpolating with the un-looped state rescues it — independent
convergent evidence for our anchor-dominant merge; their whole setting
corresponds to the untrained cell of our attribution table (17.9% hard =
our FF/untrained level), evaluated by likelihood rather than execution.
**McLeish et al. (arXiv:2511.07384)** retrofit depth-recurrence via layer
surgery + ~50B-token continued pretraining of all parameters; they name
layer choice as an open problem — §3.1 is a causal answer. Their surgery
needs a healing phase; our k=0 is exactly the base model. **Lys et al.
(arXiv:2602.14759)** loop frozen models training-free; their finding that
naive looping degrades while interpolation with the un-looped state rescues
it is independent convergent evidence for anchor-dominance, and their
setting is the untrained cell of our ladder (17.9%).
**One mechanism, three regimes.** All three works are variants of a single
design: mix the fed-back state with an anchor derived from the un-looped
computation. Lys et al.'s inference-time moving average η·h⁽⁰⁾+(1−η)·h⁽ᵗ⁾ is
an untrained anchor-mixing coefficient; our (1−α)e + α·ŝ + MLP([e;ŝ]) is its
trained analogue (fixed mix + learned correction); McLeish et al.'s
concatenated input injection is the fully learned limit, trained end-to-end.
The anchor coefficient is the stability dial of frozen-band looping: Lys
et al.'s naive-looping collapse is the zero-anchor (α→1) limit, their
regularization gains are the untrained anchored regime, and our 12B failure
at the 2B-tuned α=0.3 — with the untrained-substrate arm showing the damage
is pre-training-of-the-adapter — is the same dial mis-set at a new scale.
Stability of retrofitted recurrence appears to be governed by how strongly
the loop is anchored, across all three training budgets.
**One mechanism, three regimes.** All three works mix the fed-back state
with an anchor from the un-looped computation. Lys et al.'s moving average
η·h⁽⁰⁾+(1−η)·h⁽ᵗ⁾ is an untrained anchor coefficient; our
(1−α)e + α·ŝ + MLP is its trained analogue; McLeish et al.'s input injection
is the fully-learned limit. The 12B episode closes the loop on this
unification: the coefficient is the stability dial, naive looping is its
α→1 collapse limit, and our scale failure + state-dependent fix show the
dial must itself become a function of the state as models grow. Our stacking
results add a caution for the whole family: if retrofitted recurrence
content is amortizable (§3.3), some of the family's gains may be
reproducible by distillation without inference-time recurrence — a control
neither bracket paper runs.
Earlier lineage: Universal Transformers (adaptive depth); DEQ (fixed-point
inference); Huginn (arXiv:2502.05171) — prelude/core/coda from scratch;
Mixture-of-Recursions (arXiv:2507.10524) — learned per-token depth; Relaxed
Recursive Transformers (arXiv:2410.20672) — uptrained tied layers; Coconut —
latent CoT; pause tokens (Goyal et al.) — token-space silent compute, whose
trained-adapter variant proved a near-match for our loop on MBPP (§3.2).
Earlier lineage: Universal Transformers; DEQ; Huginn (2502.05171);
Mixture-of-Recursions (2507.10524); Relaxed Recursive Transformers
(2410.20672); Coconut; pause tokens (Goyal et al.) — whose trained variant
proved a genuine rival, not a strawman (§3.2, §3.5).
What remains distinct here: **interpretability-derived loop placement with
causal validation** (answering McLeish et al.'s open problem); **a 1.6M-param
trained merge on a fully frozen base** (between Lys et al.'s free end and
McLeish et al.'s full-retraining end, and the only one of the three where
the base model is provably untouched); **prompt-only latent planning with
bit-exact KV-cache write-in and zero decode cost**; **the attribution
ladder** (untrained / weights / pause / loop / explicit plan) — neither
bracket paper runs compute-matched token-space controls; and **difficulty-
adaptive depth via the STaR-label gate**, named as future work in both.
**Saunshi et al. (2025)** argue looped transformers trade composition
against memorization: looping buys iterative reasoning, not fact storage.
Our results reproduce this axis *within one frozen model*: k>0 moves only
the plan-dependent (compositional) slice, leaves recall-flavored MC
benchmarks flat (§3.8), and the content-injecting distill arm — not the
loop — is what nudges knowledge benchmarks up. Their looping-based
regularization (loop harder on reasoning, relax for retrieval) has an
inference-time analogue in our difficulty gate: route predicted
plan-dependent prompts to k=4 and everything else to k=0, which is the
exact base model. Retrofit looping makes the composition/memorization
trade a *per-prompt routing decision* instead of a pretraining commitment.
What remains distinct here: interpretability-derived placement with causal
validation; a fully frozen base with bit-exact k=0 and zero-decode-cost KV
write-in; the complete attribution ladder including compute-matched
token-space baselines and stacking tests; the width/depth task pattern; and the
amortizability finding itself.
## 5. Limitations
One base model family at 2B-effective scale (12B replication in progress);
two task families. **Location specificity is not yet ablated**: a
pre-registered control looping shifted/early/late/width-matched bands with
identical adapter and curriculum is queued; until it lands, the results are
formally consistent with "any wide mid-depth band works", and the lens claim
rests on discovery convenience plus mechanism verification. Hard buckets are
small (n=55 greedy / n=33 sampled) with seed spread of ±6 items; sampled
relabeling shows 97% agreement with greedy labels, and intervals accompany
all bucket cells in the final tables. The MBPP attribution grid lacks a
pause-token arm and a plan-distillation baseline (both queued) — the GSM8K
grid has the former. Easy-item perturbation tax is not eliminated (gate
preserves easy items but probe precision is 19%). Visible planning remains
stronger on absolute accuracy — the claim is cost-and-latency-shaped.
**Mixed-task training regressed both tasks**, so the current recipe yields
per-task adapters, not one general silent-planning mode; the outlook's
"installed base" framing inherits this caveat until a gate-plus-multiple-
adapters (or interference-free training) configuration is shown. MBPP
likely overlaps the base model's pretraining data; both arms share any
contamination, and memorized items land in the easy bucket, so the hard
bucket if anything over-represents genuinely novel problems — but bucket
composition is contamination-sensitive. Sensitivity to α=0.3 and band width
is unreported (the width-matched ablation arm partially addresses width).
Adapter-only training may underestimate the ceiling (band-LoRA "rung 2"
untested).
One model family (gemma-4), two scales, three task families. Hard buckets
are small (n=55/38/25); within-ladder orderings are not individually
significant, and only the pooled hard effect and the HumanEval overall gain
survive multiple-comparison scrutiny. Bucket membership derives from greedy
labeling runs (consensus-k0 robustness check moves numbers <2 points, but
both checks share the base model; an independent 12B-relabeling proxy is
running). A third architecture family was not run; rung-2 was not run at
12B. LiveCodeBench: single seed per arm, hard n=25, stdin-judged problems
only, and its newest shard (Apr 2025) is *newer than MBPP by years* but
not provably past the base model's undisclosed training cutoff — we claim
recency, not proven non-contamination. The capability panel is
MC-likelihood, not generation; its "no damage" answer does not extend to
generation quality off-distribution (LCB shows trained arms *do* hurt
there). The easy-item perturbation tax persists wherever the
gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both
arms share contamination, and memorized items land in the easy bucket, but
bucket composition is contamination-sensitive. The capability panel
(§3.8) is pending; until it lands, off-task effects of k>0 are unmeasured.
The Blocksworld-12B and GSM8K-12B failures mean the adaptive-α fix is
demonstrated on one task at one scale, not established as a general recipe.
## 6. Outlook
## 6. Conclusion
The retrofit recipe — lens-locate, anchor-merge, verifier-filtered
curriculum, gate — is scale-portable by construction: trainable mass is
independent of base size, and prompt-side loops are prefill-shaped, so their
economics *improve* with scale while serial CoT decode gets slower. The open
question that decides whether this is a curiosity or a method is whether the
effect survives scale (12B next; then a mid-size uptraining of the band
itself). If it does, "loopification" becomes a cheap post-training phase any
holder of a pretrained model can apply — a silent planning mode for the
installed base, with its latent reasoning legible to the same lens that
built it.
The experiment this program set out to run — *can an interpretability lens
tell you where to install recurrence in a frozen model, and does it work?* —
has a clean answer: yes, and the placement is causally load-bearing. The
more interesting answer is what the recurrence turned out to be: not a
reasoning engine, but a remarkably cheap way to make a frozen model amortize
its own planning into 0.03% of extra parameters, with a training-time loop
as scaffold and an inference-time loop that is optional once the content
exists. The practical recipe that survives all controls: lens-locate the
band; anchor-merge with a state-dependent coefficient; label difficulty by
STaR; distill plans if you have them, loop if you don't; gate by predicted
difficulty; keep k=0 as the exact base model. What it buys: the
plan-dependent slice at zero tokens and zero decode cost. What it does not
buy: overall accuracy beyond what matched-compute sampling already delivers.
Both halves of that sentence are the contribution.
## Appendix A: compute accounting (FLOPs / wall-clock / tokens, separated)
Let P = prompt tokens, G = generated tokens, c = FLOPs per token per full
forward pass. The band is 17 of 35 decoder layers at E2B (fraction
f≈0.486) and 10 of 48 at 12B (f≈0.208).
**Latent loop, k=4, prompt-only.** Prefill: 1 base pass + 4 band passes
over prompt positions = (1+4f)·cP ≈ **2.94·cP** at E2B (1.83× at 12B —
the overhead *shrinks* with scale because lens bands grow sublinearly).
Decode: exactly cG (looped states written to the KV cache once;
bit-exactness verified). Wall-clock: prefill is compute-bound and
position-parallel, but the k iterations are serial — prefill latency
≈2.9×, typically a small fraction of end-to-end latency for G≫0.
Visible tokens: +0.
**Best-of-3 sampling.** FLOPs: with shared prompt prefill (favorable
accounting), cP + 3·cG ≈ cP + 3cG; without sharing 3c(P+G). For MBPP
(P≈150300, G≈150220), the *extra* FLOPs vs direct (≈2cG) are of the same
order as the loop's extra (≈1.94cP) — hence "≈matched". Wall-clock: 3×G
serial bandwidth-bound decode steps (or 3 parallel decode streams at 3×
memory); strictly worse latency than the loop unless parallelized.
Visible tokens: ≈3× (two discarded candidates). Requires a verifier or
selector to pick among samples for the overall win we report (we use
any-pass, an upper bound — see §3.4 caveat).
**Budget-CoT (50-token plan).** FLOPs: ≈c(P+G+50) plus the plan tokens'
KV in context for the remainder — the *cheapest* arm in FLOPs. Wall-clock:
+50 serial decode steps before answer tokens start (worst first-token
latency). Visible tokens: +50.
Summary: no single scalar makes these three arms "equal"; the loop
dominates on tokens and decode latency, budget-CoT on FLOPs, best-of-3 on
overall accuracy. §3.4's "≈matched FLOPs" refers to the extra-FLOPs
order-of-magnitude equivalence above, not exact equality; the honest
statement is the three-way trade-off, and we report all three axes.
+131
View File
@@ -0,0 +1,131 @@
# Prototype plan: the self-paced workspace (v2)
*Drafted 2026-07-16, pre-registration-style. Goal: test whether the model
can learn to allocate workspace-loop compute ON ITS OWN — per prompt and
per generation step — rather than at a swept hyperparameter k.*
## The concept
At every step the system chooses: emit, or spend a band iteration updating
the workspace first. Make that choice a learned gate g(workspace state).
Compute becomes a decision, not a constant. Gate-bought iterations emit no
tokens, so they are exposure-safe by construction (deterministic given
state — the pause-position property).
## What already exists (de-risked ingredients)
| ingredient | evidence | where |
|---|---|---|
| static per-prompt gate (E0) | gated 52.0 overall, easy 97.5 (vs 88.5 uniform), hard 28.6; bottleneck = probe recall (18/28 tp, 95 predicted hard) | `gate_probe.py`, `eval_gated.json` |
| state-dependent control heads train | adaptive-α rescued 12B (3.8K params) | AdaptiveMergeAdapter |
| evolving state during generation | carry beats registers on GSM (hard 0→9.4) | `carry_common.py`, `eval_carry.json` |
| loop-capacity knob | rung-2 band-LoRA = best hard numbers (42.9/46.4) | `lora_band.py` |
| gates over DEPTH (complementary axis) | learnable compute envelope g[t,l] | `path_gates.py` (Nils, in progress) |
| free gate supervision | STaR difficulty labels; per-position labels derivable | prep_star/prep_mbpp |
## Experiments
### E1 — learned per-prompt halting (prompt side, MBPP)
Replace fixed k with a trained soft halting gate. Architecture: after each
iteration i, gate head h(e, ŝ_i) → p_halt,i (zero-init to fixed-k
behavior); training uses the soft mixture of iteration outputs weighted by
halting distribution (ACT-style), CE + λ·E[iterations] compute penalty;
deploy = argmax halt. Trains end-to-end, NO RL. Arms: λ ∈ {1e-3, 1e-2},
vs E0 probe-gate and uniform-k anchors.
**Pre-registered predictions:** (a) accuracy ≥ uniform k=4 overall at ≤60%
of its mean iterations; (b) easy ≥ 95% (gate protects the substrate);
(c) allocation correlates with STaR label (point-biserial r > 0.3);
(d) hard ≥ E0's 28.6% (learned gate beats frozen probe recall).
**Failure mode to watch:** gate collapse (all-0/all-1) — mitigate with
penalty warmup + entropy bonus; collapse at all λ falsifies E1.
### E2 — generation-side gating (GSM, gated carry)
Substrate: design-C carry + short verified-CoT supervision (dense targets;
harvest with "solve in ≤3 short steps", answer-verified). Gate per token
step decides whether the carry state updates through the band or passes
through: x_t = g·merge(e_t, s_{t1}) + (1g)·e_t, penalty λ·E[g].
Anchors: carry-always, carry-never (same supervision).
**Predictions:** (a) gate fires non-uniformly, concentrated near numeric/
operator tokens (measurable); (b) accuracy ≥ carry-always (gating as
protection); (c) easy-bucket damage < carry-always's (83→45 was the
unprotected number). Hard-bucket *gain* over carry-always is hoped for,
not predicted.
### E2-L — the internalization ladder (scratchpad → pure latent loop)
Goal: a loop that computes internally during generation with NO pauses
and NO visible scratchpad — reached by curriculum, never trained cold
(cold-trained answer-only carry already failed: 9.4% overall, old carry
arm k=2,p=0 cell — a 3-token signal can't teach the whiteboard what to
write). Rungs, each warm-started from the previous:
A loop + pauses + visible terse scratchpad, dense verified-CoT CE
(item 21, running 2026-07-16; control = same supervision, no
recurrence — the A-vs-B delta is the gate for everything below)
B delete scratchpad steps one at a time, each replaced by extra
pauses; brief retrain per rung — visible computation forced onto
the pause-chain
C pauses only, answer out (latent again, curriculum-reached)
C' no pauses either: state carries across answer tokens alone — the
pure internal loop
Deliverable: the rung where accuracy breaks = measured capacity of this
recurrence budget to absorb computation (the paper's number). Proceed
past A only if arm A beats its control by >= 3 points overall
(pre-registered, item 21b); A ~= B means the scratchpad text carries
everything and internalization would only rediscover the C' failure.
### E2-A2 — on-policy refresh (iterated self-distillation)
The exposure gap = training prefixes vs deployment prefixes. Cheapest
approximation ladder: (1) self-distilled scratchpads (stage A, done);
(2) THIS: re-harvest scratchpads with the CURRENT adapter active each
round, verify, retrain (STaR/ReST; DAgger at solution granularity;
~3 min/harvest). Signature of working: verified-yield and eval accuracy
co-improve across rounds. Gated on the A-vs-B verdict.
### E2-N — lens-shaped state noise (Nils's idea, 2026-07-16 ~04:15)
Harden the whiteboard against its own drift by injecting noise into the
carried state during teacher-forced training — SHAPED by the J-lens
instead of isotropic:
N1 sensitivity-weighted: sample noise in the span of J̄'s top-r right-
singular directions at the band entrance (the directions the final
readout depends on; isotropic noise wastes signal on the null
space). Cheap: jbar.pt exists; --lensnoise rank,scale flag.
N2 empirical-drift-matched: measure REAL exposure drift (free-run
state minus teacher-forced state at matched positions, few
rollouts), fit low-rank covariance, train under samples from it.
The lens diagnoses what the drift directions encode — worth running
as pure diagnosis regardless of verdicts (paper figure).
N3 concept-jitter: lens-read the carried concept (e.g. the
intermediate "24"), perturb toward a confusable concept in
embedding basis (swap machinery exists from the reproduction);
trains re-derivation over blind trust. Most ambitious.
Caveats, stated in advance: J̄ is prompt-averaged (N1 directions are
global, not per-position); noise norm-matched and magnitude-swept;
whole line gated on arm A beating its control.
### E3 — power knob (only if E1 or E2 shows clean gating)
Warm-start rung-2 band-LoRA under the gate; joint fine-tune. Question: do
gate-bought iterations do MORE per iteration with a trainable band?
Metric: the internalization count (how many scratchpad steps can be
removed post-hoc, E2 curriculum) as a function of LoRA rank.
### Lens verification (throughout — our home advantage)
J-lens reads of gated vs ungated positions: do bought iterations sharpen
task-relevant concepts at the positions where the gate fired? This is the
mechanistic check that the gate allocates *meaningfully*, not just
correlationally.
## Explicitly out of scope for the prototype
Outcome-RL training of the gate (GRPO with compute price) — stage 2, only
if E1E3 show selective gating. 12B/scale transfer. Cross-task gates.
## Budget & order
E1: 3 arms × ~75 min (Spark). E2: harvest ~30 min + 3 arms × ~90 min.
E3: +2 arms. Total ≈ 1.5 Spark-days. Runs after the lens campaign; queue
via gpuq as usual, every arm pre-registered in PROTOCOL_UNIFIED.md before
launch (items 18+).
## Kill criteria (decided in advance)
- E1 gate collapse at all λ AND E2 uniform firing → the state does not
carry usable "needs compute" signal at this scale; program stops, E0's
static-gate deployment note stands as the practical answer.
- E1 works but hard < E0 → learned gate worse than probe; ship probe-gate,
keep E2 only if its (a)/(b) hold.
+8
View File
@@ -158,3 +158,11 @@ directly to Lys's distribution-shift account and predicts the queued
auto-alignment (adaptive η, training-free) is the natural fallback if no
fixed α transfers. Worth a paragraph in Paper A (design justification +
12B analysis) and Paper B (stability mechanism).
**Outcome (2026-07-14, prediction confirmed):** the `12b_adaptive` MBPP
arm (adaptive anchoring, 2×/8× H100 fleet) recovers the substrate —
k=0/2/4 = 72.4/70.8/69.4 overall vs 72.6/54.6/43.0 under fixed α=0.3,
with hard monotone 11.4→18.2→27.3. The tolerable-loop-share account
called this shape in advance: stability restored by weakening the loop
share, hard gains preserved and k-monotone, residual easy erosion
(98.9→92.2) to be handled by the gate.
+78
View File
@@ -0,0 +1,78 @@
# The full matrix: items 131 (as of 2026-07-17 late)
Scored source of truth: PROTOCOL_UNIFIED.md. All items pre-registered before running.
## Arc I — Band-loop retrofit (MBPP; frozen E2B + merge adapter)
| # | tried | key result | control/reference | verdict |
|---|---|---|---|---|
| 14 | unified merge adapter, k-loop over prompt | MBPP hard 3.6→28.6 (k=2); GSM hard 0→6.3 | k=0 same harness | loop works on code; GSM fails from day one |
| 5 | same-size adapter, no recurrence | 17.9 hard | vs 28.6+ looped | loop > weights — recurrence load-bearing |
| 6 | mixed-task training | both tasks regressed | single-task arms | interference, no synergy |
| 7 | band-location ablation | L1430: 43.6 vs early 23.6 / shifted 21.8 | width-matched | lens's "where" confirmed; some bands structurally null (KV-share) |
| 9 | anchor sweep 13/12/11 | 34/31/25% overall, monotone collapse | L14 anchor | L14 boundary special |
| 10 | L9 anchor (full-attention layer) | catastrophic (≤21.4 hard) | anchors 1113 | lens boundary, not layer type |
## Arc II — Adapter-class factorial
| # | tried | key result | verdict |
|---|---|---|---|
| 11 | unconstrained RecurrentAdapter | ρ→4.5, easy 98→69, no depth gain | 4× params bought nothing |
| 12 | Parcae (ρ<1 certified) | robust training, saturates, easy still 71% | stability ≠ fidelity — independent dials |
| 13a | per-depth adapters (LTV) | hard content depth-stranded (17.9 @ k=2) | weight-sharing load-bearing |
| 13b | free-ACT probe | no state fixed point; no difficulty gradient | output-stable orbit; no free halting |
| 14 | tied-alpha (anchored B) | easy 93 preserved; α never moves | free B = fidelity culprit; 0.3 optimal |
| 15 | randk / noise-s₀ / h2048 / seeds | 5054 cells → seed means 3746 | lucky seeds; only seed means are levels |
| 16 | cross-task transfer | code adapter on GSM toxic (easy →2845%) | content task-local, monotone ladder |
## Arc III — Gates & the GSM boundary
| # | tried | key result | verdict |
|---|---|---|---|
| 17 | GSM-only training, best recipe | hard ≤8.7 | structural → supervision-density diagnosis |
| 18/19 | learned halting heads (E1a/b/c) | all lose to E0 frozen probe; E1c easy routing 95.9 @ 0.11 iters | classifier quality binds; hard recall regressed |
| 20 | threshold curve + oracle | oracle 59.6 @ 0.24 iters; hards depth-diverse | gate worth ~9.6 pts, unclaimed |
## Arc IV — Hybrid & internalization ladder (GSM)
| # | tried | matched | ablated/control | verdict |
|---|---|---|---|---|
| 21 | carry + dense self-distilled scratchpads | **57.4** | FF control 54.7 | 5× prior best; recurrence edge = drop bucket, p=0.0094 |
| 22 | delete steps d=1/2/3, +10 pauses each | 31.6 / 18.4 / 19.1 | base 10.9, cold 9.4 | breaks at d=1; plateau 2× cold (p=0.0025) |
| 23a | 3× training steps | val ↑ (overfit) | — | time not binding |
| 23b | 3× pauses | 29.3 | 31.6 | bandwidth not binding |
| 24 | loop-only band-LoRA r16 | stopped (fit unchanged) | — | expressivity not binding; k=0 bit-exact validated |
## Arc V — Lens/state supervision of the latent chain
| # | tried | matched | ablated | verdict |
|---|---|---|---|---|
| 25 | lens-CE: pause j ↔ deleted token j | 31.2 | — | lce 10.3→1.9 yet flat: writing ≠ computing |
| 26 | result-staging on pre-'=' spans | 24.2 / 15.6 combined | — | harmful — violates just-in-time schedule |
| 27 | zero-pause 10-iter burst + lens | 34.0 | 32.4 (p=0.45) | best nominal, ns; pause tape dead weight |
| 28 | teacher-state endpoint distillation | 30.1 | 29.7 | cos .113→.044, function absent; easy damaged |
## Arc VI — 2026-07-17 designs (Nils)
| # | tried | matched | ablated/ref | verdict |
|---|---|---|---|---|
| 29 | trajectory TF (10 waypoint transitions) | **39.1, p=0.045** | 39.1 burst-off (p=1.0) | first significant positive — a training signal, not an inference loop; +fr 30.5 (hurts); answer-only 15.6 < 19.1 |
| 30 | metacog readiness head on carried state | AUC 0.798 | FF AUC 0.791 | signal real & cheap, NOT recurrence-specific; early-stop loses; oracle +3.9 |
| 31 | synthetic KV memory (per-layer prefix) | 39.5 | 39.1 (p=1.0) | flat — gates frozen at 10 (init gradient-trap confound; 3 rerun open) |
| 32 | discrete latent chain ("latent paper": lens-snapped symbols fed back) | TF 20.3 / ST 27.0 | 39.1 | net-harmful — exposure catastrophe (TF) and quantization noise (ST) both lose to the pure analog carry; architecture tree closed |
Instrument (unnumbered): whiteboard microscopy — three specimens + carry-vs-FF divergence (probe_discount*/probe_gsm*; board artifact).
## Standing positives
MBPP loop-vs-weights gap · GSM drop-bucket reach (p=0.0094) · the 57.4 hybrid · trajectory-TF as a gradient (p=0.045) · 0.8-AUC readiness probe.
## Standing walls
Consumption (8 write-side axes + 1 read-path attempt) · internalization (4 capacity axes + 4 supervision forms) · learned gates < frozen probe.
## Open threads
Seeds for 39.1 · trajectory TF on rung A (move 57.4) · KV gate-init 3 · extension/deferral gating · g into the depth gate.
## Final architecture verdict (item 32 closes the tree)
Pauses, bursts, KV memory, analog TF chains, and discrete chains all have controlled answers.
The loop is a plan machine; tokens are the executor — they win by discreteness PLUS a verified
commitment distribution (the LM head is trained to commit; the lens readout is not).
File diff suppressed because it is too large Load Diff
+73
View File
@@ -0,0 +1,73 @@
# Final statistics pass
## Headline numbers (Wilson 95% CIs)
- **MBPP untrained merge k=4 (n=500 rerun)**: overall 0.502 [0.458, 0.546] (n=500); hard 0.200 [0.116, 0.324] (n=55)
- **MBPP trained FF k=1 (n=500 rerun)**: overall 0.536 [0.492, 0.579] (n=500); hard 0.273 [0.173, 0.402] (n=55)
- **MBPP loop s0 k=0 (base)**: overall 0.518 [0.474, 0.561] (n=500); hard 0.055 [0.019, 0.149] (n=55)
- **MBPP loop s0 k=2**: overall 0.530 [0.486, 0.573] (n=500); hard 0.309 [0.203, 0.440] (n=55)
- **MBPP loop s0 k=4**: overall 0.536 [0.492, 0.579] (n=500); hard 0.436 [0.314, 0.567] (n=55)
- **MBPP distill s1 k=1 (FF)**: overall 0.544 [0.500, 0.587] (n=500); hard 0.418 [0.297, 0.550] (n=55)
- **MBPP pause16 k=1**: overall 0.552 [0.508, 0.595] (n=500); hard 0.364 [0.249, 0.496] (n=55)
- **MBPP stack-train k=4**: overall 0.524 [0.480, 0.567] (n=500); hard 0.345 [0.234, 0.477] (n=55)
- **MBPP distill-in-loopmode k=2**: overall 0.458 [0.415, 0.502] (n=500); hard 0.200 [0.116, 0.324] (n=55)
- **Rust transfer k=0**: overall 0.591 [0.512, 0.665] (n=154); hard 0.080 [0.022, 0.250] (n=25)
- **Rust transfer k=4**: overall 0.552 [0.473, 0.628] (n=154); hard 0.240 [0.115, 0.434] (n=25)
- **HumanEval loop k=4**: overall 0.665 [0.589, 0.732] (n=164); hard 0.316 [0.191, 0.475] (n=38)
- **HumanEval distill k=1**: overall 0.646 [0.571, 0.715] (n=164); hard 0.237 [0.130, 0.392] (n=38)
- **MBPP best-of-3 (compute-matched)**: overall 0.572 [0.528, 0.615] (n=500); hard 0.327 [0.218, 0.459] (n=55)
- **MBPP budget-CoT-50**: overall 0.538 [0.494, 0.581] (n=500); hard 0.400 [0.281, 0.532] (n=55)
- **MBPP distill, 8 runs (hard)**: mean 0.457 ± 0.046 sd (range 0.400-0.545); overall mean 0.555
- **MBPP loop seeds k=4 (hard)**: mean 0.375 ± 0.055 sd (n_seeds=5)
## McNemar exact tests (paired on items)
- loop k=4 vs k=0, overall: A-only 30, B-only 39, n=500, p=0.3356 (n.s.)
- loop k=4 vs k=0, hard: A-only 1, B-only 22, n=55, p=5.722e-06 (**significant**)
- loop k=4 vs UNTRAINED merge k=4, hard (net effect): A-only 4, B-only 17, n=55, p=0.007197 (**significant**)
- loop k=4 vs UNTRAINED merge k=4, overall: A-only 29, B-only 46, n=500, p=0.06395 (n.s.)
- trained FF vs UNTRAINED merge, hard: A-only 7, B-only 11, n=55, p=0.4807 (n.s.)
- distill k=1 vs UNTRAINED merge k=4, hard: A-only 3, B-only 15, n=55, p=0.007538 (**significant**)
- distill k=1 vs loop k=4, overall: A-only 33, B-only 37, n=500, p=0.7202 (n.s.)
- distill k=1 vs loop k=4, hard: A-only 10, B-only 9, n=55, p=1 (n.s.)
- stack-train k=4 vs distill k=1, hard: A-only 10, B-only 6, n=55, p=0.4545 (n.s.)
- HumanEval loop k=4 vs k=0, overall: A-only 5, B-only 18, n=164, p=0.01062 (**significant**)
- HumanEval distill k=1 vs k=0, overall: A-only 7, B-only 17, n=164, p=0.06391 (n.s.)
- HumanEval trained vs UNTRAINED merge (k=2), overall: A-only 7, B-only 9, n=164, p=0.8036 (n.s.)
- Rust loop k=4 vs k=0, overall: A-only 11, B-only 5, n=154, p=0.2101 (n.s.)
- Rust loop k=4 vs k=0, hard: A-only 0, B-only 4, n=25, p=0.125 (n.s.)
## Token baselines (paired, per-item)
- **best-of-3 ORACLE (any-pass)**: overall 0.578 [0.534, 0.621] (n=500); hard 0.345 [0.234, 0.477] (n=55)
- **best-of-3 oracle (selector run)**: overall 0.578 [0.534, 0.621] (n=500); hard 0.345 [0.234, 0.477] (n=55)
- **best-of-3 DEPLOYABLE (logprob-selected)**: overall 0.550 [0.506, 0.593] (n=500); hard 0.273 [0.173, 0.402] (n=55)
- bo3-oracle vs loop k=4, overall: arm-only 27, bo3-only 48, p=0.0203 (**significant**)
- bo3-oracle vs loop k=4, hard: arm-only 14, bo3-only 9, p=0.4049 (n.s.)
- bo3-oracle vs distill k=1, overall: arm-only 27, bo3-only 44, p=0.05681 (n.s.)
- bo3-oracle vs distill k=1, hard: arm-only 14, bo3-only 10, p=0.5413 (n.s.)
- bo3-deployable vs loop k=4, overall: arm-only 31, bo3-only 38, p=0.4704 (n.s.)
- bo3-deployable vs loop k=4, hard: arm-only 16, bo3-only 7, p=0.09314 (n.s.)
- bo3-deployable vs distill k=1, overall: arm-only 31, bo3-only 34, p=0.8043 (n.s.)
- bo3-deployable vs distill k=1, hard: arm-only 15, bo3-only 7, p=0.1338 (n.s.)
- **budget-CoT-50 (per-item rerun)**: overall 0.538 [0.494, 0.581] (n=500); hard 0.382 [0.265, 0.514] (n=55)
- budget-CoT vs loop k=4, overall: arm-only 36, cot-only 37, p=1 (n.s.)
- budget-CoT vs loop k=4, hard: arm-only 14, cot-only 11, p=0.69 (n.s.)
- budget-CoT vs distill k=1, overall: arm-only 29, cot-only 26, p=0.7877 (n.s.)
- budget-CoT vs distill k=1, hard: arm-only 8, cot-only 6, p=0.7905 (n.s.)
## Pooled hard bucket (MBPP + HumanEval + Rust)
Paired within-item k>0 vs k=0, counts pooled across benchmarks (loop arm; distill pooled where available).
- **loop**: base 5/118 -> loop 42/118 (0.042 -> 0.356, CI [0.275, 0.446]), McNemar p=1.46e-10
- **distill**: base 3/93 -> distill 32/93 (0.032 -> 0.344, CI [0.255, 0.445]), McNemar p=2.98e-08
## Label robustness (consensus-k0 hard set)
Hard bucket redefined as: labeled hard AND k=0 fails in every seed's own eval run (removes single-greedy-run selection noise).
- consensus hard set: 52 of 55 labeled-hard items
- MBPP loop s0 k=4: labeled-hard 0.436 -> consensus-hard 0.423 [0.299, 0.558] (n=52)
- MBPP distill s1 k=1 (FF): labeled-hard 0.418 -> consensus-hard 0.404 [0.282, 0.539] (n=52)
- MBPP stack-train k=4: labeled-hard 0.345 -> consensus-hard 0.346 [0.232, 0.482] (n=52)
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+47
View File
@@ -0,0 +1,47 @@
{
"tag": "rec16_e400_PARTIAL",
"note": "k=16,32 cancelled by decision after k<=8 showed prediction (a); rho(A) from checkpoints",
"rho_trajectory": {
"e100": 3.357,
"e200": 4.31,
"e300": 4.456,
"e400": 4.525,
"e500": 4.492,
"e600": 4.465
},
"ks": {
"0": {
"acc": 0.488,
"by_label": {
"easy": 0.984,
"hard": 0.036,
"drop": 0.01
}
},
"2": {
"acc": 0.408,
"by_label": {
"easy": 0.697,
"hard": 0.357,
"drop": 0.07
}
},
"4": {
"acc": 0.412,
"by_label": {
"easy": 0.697,
"hard": 0.393,
"drop": 0.07
}
},
"8": {
"acc": 0.408,
"by_label": {
"easy": 0.689,
"hard": 0.429,
"drop": 0.06
}
}
},
"n": 250
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+8
View File
@@ -0,0 +1,8 @@
{
"tag": "base_k0",
"k": 0,
"arc": 0.36,
"winogrande": 0.55875,
"hellaswag": 0.5225,
"mmlu": 0.30125
}
+8
View File
@@ -0,0 +1,8 @@
{
"tag": "distill_ff",
"k": 1,
"arc": 0.4175,
"winogrande": 0.56625,
"hellaswag": 0.57,
"mmlu": 0.3175
}
+8
View File
@@ -0,0 +1,8 @@
{
"tag": "loop_k2",
"k": 2,
"arc": 0.36125,
"winogrande": 0.5525,
"hellaswag": 0.49625,
"mmlu": 0.3125
}
+26
View File
@@ -0,0 +1,26 @@
{
"0": {
"acc": 0.5,
"by_label": {
"easy": 0.9672131147540983,
"hard": 0.17857142857142858,
"drop": 0.02
}
},
"2": {
"acc": 0.504,
"by_label": {
"easy": 0.9098360655737705,
"hard": 0.35714285714285715,
"drop": 0.05
}
},
"4": {
"acc": 0.512,
"by_label": {
"easy": 0.9098360655737705,
"hard": 0.42857142857142855,
"drop": 0.05
}
}
}
+26
View File
@@ -0,0 +1,26 @@
{
"0": {
"acc": 0.496,
"by_label": {
"easy": 0.9672131147540983,
"hard": 0.10714285714285714,
"drop": 0.03
}
},
"2": {
"acc": 0.516,
"by_label": {
"easy": 0.9344262295081968,
"hard": 0.39285714285714285,
"drop": 0.04
}
},
"4": {
"acc": 0.524,
"by_label": {
"easy": 0.9098360655737705,
"hard": 0.4642857142857143,
"drop": 0.07
}
}
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 202 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 84 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 104 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 171 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 66 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 80 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 43 KiB

File diff suppressed because it is too large Load Diff
+93
View File
@@ -0,0 +1,93 @@
{
"curve": [
{
"theta": 0.3,
"overall": 0.488,
"by_label": {
"easy": 0.9672131147540983,
"hard": 0.07142857142857142,
"drop": 0.02
},
"ek": 0.432
},
{
"theta": 0.5,
"overall": 0.5,
"by_label": {
"easy": 0.9590163934426229,
"hard": 0.17857142857142858,
"drop": 0.03
},
"ek": 0.74
},
{
"theta": 0.7,
"overall": 0.496,
"by_label": {
"easy": 0.9426229508196722,
"hard": 0.17857142857142858,
"drop": 0.04
},
"ek": 1.136
},
{
"theta": 0.8,
"overall": 0.508,
"by_label": {
"easy": 0.9344262295081968,
"hard": 0.2857142857142857,
"drop": 0.05
},
"ek": 1.388
},
{
"theta": 0.9,
"overall": 0.496,
"by_label": {
"easy": 0.8852459016393442,
"hard": 0.39285714285714285,
"drop": 0.05
},
"ek": 1.76
},
{
"theta": 0.95,
"overall": 0.492,
"by_label": {
"easy": 0.8770491803278688,
"hard": 0.42857142857142855,
"drop": 0.04
},
"ek": 2.084
},
{
"theta": 0.98,
"overall": 0.504,
"by_label": {
"easy": 0.8770491803278688,
"hard": 0.4642857142857143,
"drop": 0.06
},
"ek": 2.412
},
{
"theta": 0.99,
"overall": 0.508,
"by_label": {
"easy": 0.8770491803278688,
"hard": 0.5,
"drop": 0.06
},
"ek": 2.628
}
],
"oracle": {
"overall": 0.596,
"by_label": {
"easy": 1.0,
"hard": 0.6428571428571429,
"drop": 0.09
},
"ek": 0.236
}
}
File diff suppressed because it is too large Load Diff
+152
View File
@@ -0,0 +1,152 @@
{
"abc382_a": true,
"abc382_b": true,
"abc382_c": false,
"abc382_d": true,
"abc382_f": false,
"abc382_g": false,
"abc383_a": true,
"abc383_b": true,
"abc383_c": false,
"abc383_d": true,
"abc383_e": false,
"abc384_a": true,
"abc384_b": true,
"abc384_c": true,
"abc384_d": false,
"abc384_e": false,
"abc384_f": false,
"abc384_g": false,
"abc385_a": true,
"abc385_b": true,
"abc385_c": false,
"abc385_d": false,
"abc385_e": false,
"abc385_f": false,
"abc386_a": false,
"abc386_b": true,
"abc386_c": false,
"abc386_d": false,
"abc386_e": false,
"abc386_f": false,
"abc387_a": true,
"abc387_b": true,
"abc387_c": false,
"abc387_f": false,
"abc388_a": true,
"abc388_b": true,
"abc388_c": false,
"abc388_d": false,
"abc388_e": false,
"abc388_f": false,
"abc388_g": false,
"abc389_a": true,
"abc389_b": true,
"abc389_d": false,
"abc389_e": false,
"abc389_f": false,
"abc389_g": false,
"abc390_a": true,
"abc390_b": true,
"abc390_c": false,
"abc390_d": false,
"abc390_e": false,
"abc390_f": false,
"abc390_g": false,
"abc391_a": true,
"abc391_b": true,
"abc391_d": false,
"abc391_e": false,
"abc391_f": false,
"abc391_g": false,
"abc392_a": true,
"abc392_b": true,
"abc392_c": false,
"abc392_d": false,
"abc392_f": true,
"abc392_g": false,
"abc393_a": true,
"abc393_b": true,
"abc393_d": false,
"abc393_e": true,
"abc393_f": false,
"abc394_a": true,
"abc394_b": true,
"abc394_c": true,
"abc394_d": true,
"abc394_e": false,
"abc394_f": false,
"abc394_g": false,
"abc395_a": true,
"abc395_b": true,
"abc395_c": true,
"abc395_e": true,
"abc395_f": false,
"abc396_a": true,
"abc396_b": true,
"abc396_c": false,
"abc396_d": true,
"abc396_e": false,
"abc396_f": false,
"abc396_g": false,
"abc397_a": true,
"abc397_b": false,
"abc397_c": true,
"abc397_d": false,
"abc397_e": false,
"abc397_f": false,
"abc397_g": false,
"abc398_a": true,
"abc398_b": false,
"abc398_c": true,
"abc398_d": false,
"abc398_f": true,
"abc398_g": false,
"abc399_a": true,
"abc399_b": true,
"abc399_c": true,
"abc399_d": false,
"abc399_e": false,
"abc399_f": false,
"abc400_a": true,
"abc400_b": true,
"abc400_c": false,
"abc400_d": false,
"abc400_e": false,
"abc400_g": false,
"arc188_a": false,
"arc188_b": false,
"arc188_c": false,
"arc188_d": false,
"arc189_a": false,
"arc189_b": false,
"arc189_c": false,
"arc189_d": false,
"arc190_a": false,
"arc190_c": false,
"arc190_d": false,
"arc191_a": false,
"arc191_c": false,
"arc191_d": true,
"arc192_a": false,
"arc192_b": false,
"arc192_d": false,
"arc192_e": false,
"arc193_a": false,
"arc193_b": false,
"arc193_d": false,
"arc194_a": false,
"arc194_b": false,
"arc194_c": false,
"arc194_d": false,
"arc194_e": false,
"arc195_a": false,
"arc195_b": false,
"arc195_c": false,
"arc195_d": false,
"arc195_e": false,
"arc196_a": false,
"arc196_b": false,
"arc196_c": false,
"arc196_d": false
}
+1
View File
@@ -0,0 +1 @@
{"carry": 0.7981836131859141, "ff": 0.791359567919433}
File diff suppressed because one or more lines are too long
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+561
View File
@@ -0,0 +1,561 @@
{
"MBPP untrained merge k=4 (n=500 rerun)": {
"overall": {
"acc": 0.502,
"n": 500,
"ci": [
0.458,
0.546
]
},
"hard": {
"acc": 0.2,
"n": 55,
"ci": [
0.116,
0.324
]
}
},
"MBPP trained FF k=1 (n=500 rerun)": {
"overall": {
"acc": 0.536,
"n": 500,
"ci": [
0.492,
0.579
]
},
"hard": {
"acc": 0.2727272727272727,
"n": 55,
"ci": [
0.173,
0.402
]
}
},
"MBPP loop s0 k=0 (base)": {
"overall": {
"acc": 0.518,
"n": 500,
"ci": [
0.474,
0.561
]
},
"hard": {
"acc": 0.05454545454545454,
"n": 55,
"ci": [
0.019,
0.149
]
}
},
"MBPP loop s0 k=2": {
"overall": {
"acc": 0.53,
"n": 500,
"ci": [
0.486,
0.573
]
},
"hard": {
"acc": 0.3090909090909091,
"n": 55,
"ci": [
0.203,
0.44
]
}
},
"MBPP loop s0 k=4": {
"overall": {
"acc": 0.536,
"n": 500,
"ci": [
0.492,
0.579
]
},
"hard": {
"acc": 0.43636363636363634,
"n": 55,
"ci": [
0.314,
0.567
]
}
},
"MBPP distill s1 k=1 (FF)": {
"overall": {
"acc": 0.544,
"n": 500,
"ci": [
0.5,
0.587
]
},
"hard": {
"acc": 0.41818181818181815,
"n": 55,
"ci": [
0.297,
0.55
]
}
},
"MBPP pause16 k=1": {
"overall": {
"acc": 0.552,
"n": 500,
"ci": [
0.508,
0.595
]
},
"hard": {
"acc": 0.36363636363636365,
"n": 55,
"ci": [
0.249,
0.496
]
}
},
"MBPP stack-train k=4": {
"overall": {
"acc": 0.524,
"n": 500,
"ci": [
0.48,
0.567
]
},
"hard": {
"acc": 0.34545454545454546,
"n": 55,
"ci": [
0.234,
0.477
]
}
},
"MBPP distill-in-loopmode k=2": {
"overall": {
"acc": 0.458,
"n": 500,
"ci": [
0.415,
0.502
]
},
"hard": {
"acc": 0.2,
"n": 55,
"ci": [
0.116,
0.324
]
}
},
"Rust transfer k=0": {
"overall": {
"acc": 0.5909090909090909,
"n": 154,
"ci": [
0.512,
0.665
]
},
"hard": {
"acc": 0.08,
"n": 25,
"ci": [
0.022,
0.25
]
}
},
"Rust transfer k=4": {
"overall": {
"acc": 0.551948051948052,
"n": 154,
"ci": [
0.473,
0.628
]
},
"hard": {
"acc": 0.24,
"n": 25,
"ci": [
0.115,
0.434
]
}
},
"HumanEval loop k=4": {
"overall": {
"acc": 0.6646341463414634,
"n": 164,
"ci": [
0.589,
0.732
]
},
"hard": {
"acc": 0.3157894736842105,
"n": 38,
"ci": [
0.191,
0.475
]
}
},
"HumanEval distill k=1": {
"overall": {
"acc": 0.6463414634146342,
"n": 164,
"ci": [
0.571,
0.715
]
},
"hard": {
"acc": 0.23684210526315788,
"n": 38,
"ci": [
0.13,
0.392
]
}
},
"MBPP best-of-3 (compute-matched)": {
"overall": {
"acc": 0.572,
"n": 500,
"ci": [
0.528,
0.615
]
},
"hard": {
"acc": 0.32727272727272727,
"n": 55,
"ci": [
0.218,
0.459
]
}
},
"MBPP budget-CoT-50": {
"overall": {
"acc": 0.538,
"n": 500,
"ci": [
0.494,
0.581
]
},
"hard": {
"acc": 0.4,
"n": 55,
"ci": [
0.281,
0.532
]
}
},
"distill_seed_spread": {
"hard": [
0.4909090909090909,
0.41818181818181815,
0.5454545454545454,
0.43636363636363634,
0.45454545454545453,
0.4727272727272727,
0.43636363636363634,
0.4
],
"overall": [
0.55,
0.544,
0.562,
0.546,
0.566,
0.562,
0.564,
0.55
]
},
"loop_seed_spread_hard": [
0.43636363636363634,
0.41818181818181815,
0.3090909090909091,
0.32727272727272727,
0.38181818181818183
],
"mcnemar": {
"loop k=4 vs k=0, overall": {
"n": 500,
"a_only": 30,
"b_only": 39,
"p": 0.33555761823401514
},
"loop k=4 vs k=0, hard": {
"n": 55,
"a_only": 1,
"b_only": 22,
"p": 5.7220458984375e-06
},
"loop k=4 vs UNTRAINED merge k=4, hard (net effect)": {
"n": 55,
"a_only": 4,
"b_only": 17,
"p": 0.007197380065917969
},
"loop k=4 vs UNTRAINED merge k=4, overall": {
"n": 500,
"a_only": 29,
"b_only": 46,
"p": 0.06394991646706696
},
"trained FF vs UNTRAINED merge, hard": {
"n": 55,
"a_only": 7,
"b_only": 11,
"p": 0.480682373046875
},
"distill k=1 vs UNTRAINED merge k=4, hard": {
"n": 55,
"a_only": 3,
"b_only": 15,
"p": 0.007537841796875
},
"distill k=1 vs loop k=4, overall": {
"n": 500,
"a_only": 33,
"b_only": 37,
"p": 0.7202027723528613
},
"distill k=1 vs loop k=4, hard": {
"n": 55,
"a_only": 10,
"b_only": 9,
"p": 1.0
},
"stack-train k=4 vs distill k=1, hard": {
"n": 55,
"a_only": 10,
"b_only": 6,
"p": 0.454498291015625
},
"HumanEval loop k=4 vs k=0, overall": {
"n": 164,
"a_only": 5,
"b_only": 18,
"p": 0.010622024536132812
},
"HumanEval distill k=1 vs k=0, overall": {
"n": 164,
"a_only": 7,
"b_only": 17,
"p": 0.06391465663909912
},
"HumanEval trained vs UNTRAINED merge (k=2), overall": {
"n": 164,
"a_only": 7,
"b_only": 9,
"p": 0.803619384765625
},
"Rust loop k=4 vs k=0, overall": {
"n": 154,
"a_only": 11,
"b_only": 5,
"p": 0.210113525390625
},
"Rust loop k=4 vs k=0, hard": {
"n": 25,
"a_only": 0,
"b_only": 4,
"p": 0.125
},
"bo3-oracle vs loop k=4, overall": {
"n": 500,
"a_only": 27,
"b_only": 48,
"p": 0.020298406990107445
},
"bo3-oracle vs loop k=4, hard": {
"n": 55,
"a_only": 14,
"b_only": 9,
"p": 0.4048728942871094
},
"bo3-oracle vs distill k=1, overall": {
"n": 500,
"a_only": 27,
"b_only": 44,
"p": 0.056814677932839015
},
"bo3-oracle vs distill k=1, hard": {
"n": 55,
"a_only": 14,
"b_only": 10,
"p": 0.5412561893463135
},
"bo3-deployable vs loop k=4, overall": {
"n": 500,
"a_only": 31,
"b_only": 38,
"p": 0.4703685318581444
},
"bo3-deployable vs loop k=4, hard": {
"n": 55,
"a_only": 16,
"b_only": 7,
"p": 0.0931396484375
},
"bo3-deployable vs distill k=1, overall": {
"n": 500,
"a_only": 31,
"b_only": 34,
"p": 0.8043170001933986
},
"bo3-deployable vs distill k=1, hard": {
"n": 55,
"a_only": 15,
"b_only": 7,
"p": 0.13380050659179688
},
"budget-cot vs loop k=4, overall": {
"n": 500,
"a_only": 36,
"b_only": 37,
"p": 1.0
},
"budget-cot vs loop k=4, hard": {
"n": 55,
"a_only": 14,
"b_only": 11,
"p": 0.6900379657745361
},
"budget-cot vs distill k=1, overall": {
"n": 500,
"a_only": 29,
"b_only": 26,
"p": 0.7877061896700435
},
"budget-cot vs distill k=1, hard": {
"n": 55,
"a_only": 8,
"b_only": 6,
"p": 0.79052734375
}
},
"best-of-3 ORACLE (any-pass)": {
"overall": {
"acc": 0.578,
"n": 500,
"ci": [
0.534,
0.621
]
},
"hard": {
"acc": 0.34545454545454546,
"n": 55,
"ci": [
0.234,
0.477
]
}
},
"best-of-3 oracle (selector run)": {
"overall": {
"acc": 0.578,
"n": 500,
"ci": [
0.534,
0.621
]
},
"hard": {
"acc": 0.34545454545454546,
"n": 55,
"ci": [
0.234,
0.477
]
}
},
"best-of-3 DEPLOYABLE (logprob-selected)": {
"overall": {
"acc": 0.55,
"n": 500,
"ci": [
0.506,
0.593
]
},
"hard": {
"acc": 0.2727272727272727,
"n": 55,
"ci": [
0.173,
0.402
]
}
},
"pooled_hard_loop": {
"base": 5,
"arm": 42,
"n": 118,
"mcnemar": {
"n": 118,
"a_only": 1,
"b_only": 38,
"p": 1.4551915228366852e-10
}
},
"pooled_hard_distill": {
"base": 3,
"arm": 32,
"n": 93,
"mcnemar": {
"n": 93,
"a_only": 1,
"b_only": 30,
"p": 2.9802322387695312e-08
}
},
"consensus_hard": {
"MBPP loop s0 k=4": {
"acc": 0.4230769230769231,
"n": 52,
"ci": [
0.299,
0.558
]
},
"MBPP distill s1 k=1 (FF)": {
"acc": 0.40384615384615385,
"n": 52,
"ci": [
0.282,
0.539
]
},
"MBPP stack-train k=4": {
"acc": 0.34615384615384615,
"n": 52,
"ci": [
0.232,
0.482
]
}
}
}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long

Some files were not shown because too many files have changed in this diff Show More