Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f1bd5daf62 | ||
|
|
21b605c13a | ||
|
|
f0b7942c7a | ||
|
|
3db6dd8fef | ||
|
|
7ac7f9651a | ||
|
|
96afac769e | ||
|
|
64a20750bd | ||
|
|
b18ffbd739 | ||
|
|
11604d0949 | ||
|
|
c7167d512f | ||
|
|
3e02c6e24e | ||
|
|
add87cdd72 | ||
|
|
f1e6a628eb | ||
|
|
deae500ad0 | ||
|
|
1af6cb5e45 | ||
|
|
67f9c3091c | ||
|
|
6542f80339 | ||
|
|
bbf8a02c6b | ||
|
|
62ce30e557 | ||
|
|
44e814e78c | ||
|
|
0cffce876b | ||
|
|
64abfe66f3 | ||
|
|
f31e6e7e30 | ||
|
|
fe89820e64 | ||
|
|
7a141948fa | ||
|
|
46b5e1b802 | ||
|
|
f143c67b3d | ||
|
|
001a4ca5f3 | ||
|
|
7d60f37cf1 | ||
|
|
4e744e7d94 | ||
|
|
20f7014f5b | ||
|
|
272a7b1d1f | ||
|
|
34497b8835 | ||
|
|
051e90e805 | ||
|
|
1618c206ae | ||
|
|
4b9e168838 | ||
|
|
74f04d124f | ||
|
|
21908a0e33 | ||
|
|
a643b4a146 | ||
|
|
332f7f3622 | ||
|
|
dec27a4017 | ||
|
|
550d5cbe56 | ||
|
|
684e793eef | ||
|
|
f2a48ce338 | ||
|
|
dafbd8fd25 | ||
|
|
ba95d2aa3d | ||
|
|
357a5add06 | ||
|
|
5f58bc8aed | ||
|
|
600f965247 | ||
|
|
dc43769118 | ||
|
|
8187b79b50 | ||
|
|
3a0a6a96ae | ||
|
|
5d681a57b5 | ||
|
|
8ebb326745 | ||
|
|
cb0f1d950c | ||
|
|
2cd1f37f5f | ||
|
|
ed5455277b | ||
|
|
c97d9d9cda | ||
|
|
d81afef488 | ||
|
|
1e6473e1aa | ||
|
|
0f386ea29d | ||
|
|
b6d54c2726 | ||
|
|
b17d7faea0 | ||
|
|
873dc351a0 | ||
|
|
cf9e38f858 | ||
|
|
713fbdce76 | ||
|
|
f0e524cf68 | ||
|
|
ebf0eb8f49 | ||
|
|
b013ca27f1 | ||
|
|
4455d4bb9d | ||
|
|
13643fa504 | ||
|
|
fb5fb13806 | ||
|
|
0a5cd80dc1 | ||
|
|
8981cda31d | ||
|
|
b53538e82c | ||
|
|
82cc1b85b8 | ||
|
|
0e6d6adcb9 | ||
|
|
d39e02f4b9 | ||
|
|
16f4a057aa | ||
|
|
f370883e23 | ||
|
|
1076d57d96 | ||
|
|
3477fb5da0 | ||
|
|
d00f120f27 | ||
|
|
f254b20b62 | ||
|
|
d2da8044c3 | ||
|
|
0fd93cb328 | ||
|
|
1d4cffd6ad | ||
|
|
caef95f10f | ||
|
|
2ce646d175 | ||
|
|
e5fc453005 | ||
|
|
378bb36b8c | ||
|
|
7821f304f8 | ||
|
|
dcd0b0b2dc | ||
|
|
d1bafc2d4d | ||
|
|
01fd272028 | ||
|
|
87308cfb0b | ||
|
|
55665ce251 | ||
|
|
6a954bd3e4 | ||
|
|
b461970c9a | ||
|
|
aa3e11fb20 | ||
|
|
f2cbce5a03 | ||
|
|
e152a91c05 | ||
|
|
cd8e94dda5 | ||
|
|
9329f22a63 | ||
|
|
c718c636a5 | ||
|
|
d2d372c909 | ||
|
|
12ec01becc | ||
|
|
95be35d20f | ||
|
|
32728e7b07 | ||
|
|
471cf336d0 | ||
|
|
0252641701 | ||
|
|
5efe5b790a | ||
|
|
1460cf241f | ||
|
|
8fe8e34b47 | ||
|
|
507c170744 |
@@ -0,0 +1,98 @@
|
||||
# Session handoff — 2026-07-16 ~09:45
|
||||
|
||||
Previous session ended because the auto-mode permission classifier entered
|
||||
a persistent session-level lockout (blocked nearly all Bash/Write on
|
||||
"earlier conversation content" grounds, ~7h intermittent). Everything
|
||||
below is current as of ~09:45. Repo knowledge map: `INDEX.md`. Scored
|
||||
experiment record: `results-loop/PROTOCOL_UNIFIED.md` (items 1–21, all
|
||||
scored except pending stats below). Research plan: `PLAN_SELFPACED.md`.
|
||||
|
||||
## RUNNING RIGHT NOW — do not disturb, monitor these
|
||||
|
||||
### Rented H100 node ("node 1")
|
||||
- **ssh -p 52525 root@185.151.171.35** (vast.ai; greets with a banner —
|
||||
filter it). Costs money while it runs.
|
||||
- **`/root/pipeline5.log`**: lens-mapping pipeline (setsid, survives ssh
|
||||
drops). Sequence: 31B jbar DONE (saved+synced) → 31B exp4 scan DONE as
|
||||
a printed TABLE (its `regimes.pt` save crashed — known bug, see
|
||||
fix-list) → 31B weights cleaned → **26B-A4B jacobian scan RUNNING**
|
||||
(started ~07:40, ~40-90s/prompt × 256; expect done late morning).
|
||||
After it: 26B exp4 scan (same save bug will strike — harmless, the
|
||||
table survives in the log), CLEANED, PIPELINE-DONE.
|
||||
- **`/root/node_sync.sh` sidecar** (also setsid): rclone-pushes
|
||||
`/workspace/jspace/results/` (jbar_*.pt, exp4_*.log) and pipeline5.log
|
||||
to `jspace:jspace/results-lens/` every 120s. The node HAS rclone
|
||||
credentials (Nils placed them himself — never copy credentials to
|
||||
nodes from automation; hard-blocked + it's his call).
|
||||
- When PIPELINE-DONE: verify `results-lens/` has jbar_31b.pt,
|
||||
jbar_26b_a4b.pt, exp4_31b.log, exp4_26b_a4b.log — then the node can be
|
||||
destroyed (tell Nils; it's his dashboard).
|
||||
- Storage lore: /workspace = disk (survives stop, dies on destroy);
|
||||
/dev/shm = RAM (dies on stop). See memory/vast-node-storage.md.
|
||||
|
||||
### Spark (this machine)
|
||||
- gpuq worker in tmux session `gpuq_spark:gpuq0`, queue EMPTY, healthy.
|
||||
Worker code now has: data contract (`# gpuq-in:`/`# gpuq-out:` job
|
||||
comments → auto-sync inputs/outputs with the bucket, outputs every
|
||||
120s during the job) and graceful drain (`touch ~/gpuq/spark-gpu0/STOP`
|
||||
— NEVER kill the tmux session mid-job; that cost 3.5h once).
|
||||
- Thermal: box hard-froze 4× on 2026-07-15 under sustained load —
|
||||
cabinet now open, stable since. If it freezes: reboot, restart worker
|
||||
(`tmux new-session -d -s gpuq_spark -n gpuq0 'bash
|
||||
~/jspace/scripts/gpuq_worker.sh spark-gpu0 0'`). NO @reboot cron
|
||||
(Nils vetoed — crash-loop risk).
|
||||
|
||||
## IMMEDIATE PENDING (blocked by the classifier, run first)
|
||||
|
||||
1. `.venv/bin/python scripts/mcnemar_carrycot.py` — the paired p-values
|
||||
for last night's headline result, especially the drop-bucket test.
|
||||
THE morning number; item 21's outcome note says "pending".
|
||||
2. Commit everything uncommitted:
|
||||
`git add -A scripts/ PLAN_SELFPACED.md HANDOFF.md && git add -f
|
||||
results-loop/PROTOCOL_UNIFIED.md results-loop/eval_gsm_carrycot*.json
|
||||
results-loop/gsm_cot_data.json && git commit && git push origin main`
|
||||
(uncommitted: mcnemar script, lens-noise trainer changes
|
||||
[train_carry_cot.py --lensnoise, UNVERIFIED — parse-check first],
|
||||
plan E2-L/E2-A2/E2-N sections, protocol item-21 scoring, this file).
|
||||
3. Parse the 31B (and later 26B) regime tables from exp4_*.log into
|
||||
regimes_*.pt — copy the pattern in `scripts/parse_e4b_regimes.py`
|
||||
(adjust layer count: 31B=60, 26B=48? read from log). Then fill
|
||||
REGIMES.json rows (results/REGIMES.json) and re-render
|
||||
`scripts/fig_regimes.py` (5-scale figure, panels auto-fill).
|
||||
4. Patch `scripts/exp4_regimes.py` to take an explicit output path
|
||||
(torch.save to CWD-relative "results/" has now crashed 3 scans).
|
||||
|
||||
## LAST NIGHT'S RESULTS (already scored in the protocol)
|
||||
|
||||
- **Item 21, the headline**: GSM8K carry-cot (dense self-distilled terse
|
||||
scratchpads through the carry whiteboard): **57.4%** best cell vs
|
||||
12.1% all-time prior best; base 10.9. Control (same supervision, no
|
||||
recurrence): 54.7 best. Whiteboard's specific edge: DROP items
|
||||
(+9/+14 across cells) — reach into problems unreachable at labeling.
|
||||
If McNemar confirms → gates open for: stage B internalization ladder
|
||||
(PLAN E2-L), A2 on-policy refresh (E2-A2), lens-shaped noise (E2-N,
|
||||
Nils's idea, N1 already implemented as --lensnoise, unverified).
|
||||
- **Item 20**: gate threshold curve — E0's frozen logistic probe
|
||||
dominates every learned gate; ORACLE gate = 59.6% overall at 0.24
|
||||
mean iters (hard items depth-diverse: 18/28 solvable at some k, ≤13
|
||||
at any single k). Gate program continues; binding constraint =
|
||||
classifier quality on the pre-loop state.
|
||||
- **E4B lens surprise** (REGIMES.json, provisional): no E2B-style
|
||||
workspace signature — sensor L11-22, motor from L23, persistence bump
|
||||
ABSENT; candidate band sits inside the KV-shared zone. The elastic
|
||||
pair (E2B⊂E4B) reorganized. Needs ignition cross-check + jbar top-up
|
||||
before strong claims.
|
||||
|
||||
## MORNING DECISION QUEUE (Nils decides, one submit each)
|
||||
|
||||
With McNemar in hand: which of stage B (internalization) / A2
|
||||
(re-harvest) / E2-N1 (lens-noise, verify parse first) gets the Spark.
|
||||
All pre-registered or planned in PLAN_SELFPACED.md; job template
|
||||
pattern: scripts/jobs/zzz_l_gsm_e2a.sh (uses the data contract).
|
||||
|
||||
## OPERATIONAL CAUTIONS (paid for in blood, see LESSONS.md 1-12 + memory)
|
||||
|
||||
- pkill -f self-match; k=0 sanity row in every eval; e400 pre-commit;
|
||||
seeds before believing single cells (noise-s0 taught this twice);
|
||||
monitors: grep patterns must match eval_loop.py's "acc=" (GSM) vs
|
||||
"pass@1=" (MBPP); artifacts leave nodes within one sync cycle.
|
||||
@@ -0,0 +1,84 @@
|
||||
# Where to find what
|
||||
|
||||
Two projects share this repo: the **J-lens reproduction** (does the 2026
|
||||
workspace paper replicate on gemma-4-E2B?) and the **workspace-looping
|
||||
investigation** that grew out of it (retrofit recurrence onto the lens-found
|
||||
band; what does it actually buy?). The second is the active one.
|
||||
|
||||
## The claims and their evidence
|
||||
|
||||
| you want | look in |
|
||||
|---|---|
|
||||
| Current claims, all numbers, figures | `PAPER.md` (source of truth; the .pdf snapshots lag it) |
|
||||
| What worked / what failed / design rules / ops pitfalls | `LESSONS.md` — read before running anything on this hardware |
|
||||
| Pre-registrations + scored outcomes (17 items, incl. refutations) | `results-loop/PROTOCOL_UNIFIED.md` — the methods backbone; every claim in PAPER.md §3 traces to an item here |
|
||||
| Significance tests behind any claimed number | `results-loop/STATS.md` |
|
||||
| The v2 prototype plan (self-paced workspace: learned gating) | `PLAN_SELFPACED.md` |
|
||||
| Lab-notebook narrative of the looping investigation | `WORKSPACE_LOOPING.md` (superseded where it disagrees with PAPER.md) |
|
||||
| Base-reproduction results (lens replication itself) | `RESULTS.md`, `README.md` |
|
||||
| Per-model lens maps: workspace bands, KV-share boundaries, pinned revisions | `results/REGIMES.json` (canonical registry) + `results/jbar*.pt` (raw J̄) + `results/exp4*.log` (regime scans) |
|
||||
|
||||
## Code (`scripts/`, `jlens/`)
|
||||
|
||||
- `jlens/core.py` — the lens: model loading (`JLENS_MODEL` env), J̄ readouts.
|
||||
- `scripts/loop_common.py` — everything band-looping: `BandLooper`
|
||||
(capture/re-run machinery, KV-cache-safe), `generate_frozen_prompt`
|
||||
(the ≥3.5× deploy path), and every adapter variant from the regime sweep
|
||||
(`MergeAdapter` ★, `AdaptiveMergeAdapter`, `RecurrentAdapter`,
|
||||
`ParcaeAdapter`, `NoisyMergeAdapter`, `TiedAlphaAdapter`,
|
||||
`PerDepthAdapter`). Band via `JLENS_BAND` env (default E2B 14,30).
|
||||
- Trainers: `train_merge_code.py` (MBPP; all regime flags live here),
|
||||
`train_merge.py` (GSM, old full-position regime — historic),
|
||||
`train_merge_unified.py` (multi-task, hardened protocol),
|
||||
`train_merge_bw.py` (Blocksworld), `train_distill*.py` (plan distillation).
|
||||
- Evals: `eval_loop_code.py` (MBPP pass@1 vs k; per-item logs; `--halt`),
|
||||
`eval_loop.py` (GSM), `eval_bw.py`, `eval_humaneval.py`, `eval_rust.py`,
|
||||
`eval_lcb.py`, `eval_mc_panel.py`, plus `prep_*.py` (STaR labeling).
|
||||
- Figures: `fig_*.py` regenerate the canonical PNGs from the JSONs.
|
||||
- Infra: `gpuq_*.sh` + `GPUQ.md` (bucket-backed GPU job queue),
|
||||
`node_setup.sh` (vast.ai bootstrap; pins model revisions — see LESSONS #12),
|
||||
`vast-ai-notes.md`.
|
||||
|
||||
## Results directories — including the honest mess
|
||||
|
||||
- `results-loop/` — **the looping project's data**: 84 `eval_*.json`
|
||||
(tag suffixes: `_s<seed>`, `_rec16`/`_parcae16` recurrent arms, `_pd4`
|
||||
per-depth, `_ta` tied-alpha, `_rk16` random-depth, `_ns` noise-s₀,
|
||||
`_h2048` capacity, `code2gsm_*` transfer; `per_item` only in files from
|
||||
Jul 14 onward), adapter checkpoints (`adapter_*.pt`, e400 = the
|
||||
pre-committed eval checkpoint), canonical figures (`fig_kcurves.png`
|
||||
design-space grid, `fig_phase.png` two-dials diagram, `fig_loop_vs_ff.png`
|
||||
recurrence-vs-distill ladder, `fig_placement/transfer/scale.png`),
|
||||
and `chain*.log` — autonomous-session logs, archaeology only.
|
||||
- `results/` — lens reproduction outputs + the cross-model registry
|
||||
(`REGIMES.json`).
|
||||
- `results-band-*/` — one directory per entrance-placement arm of the
|
||||
placement sweep (L2–L24 entrances); summarized in PAPER fig_placement;
|
||||
kept for per-item audit.
|
||||
- `results-12b/`, `results-loop-12b/` — 12B lens map and looping arms.
|
||||
- `results-tap23/`, `results-tap34/`, `results-kvtest/`, `results-combo/`,
|
||||
`results-panel/`, `results-distill-s7/` — single-question side arms
|
||||
(exit-tap sweep, KV nulling check, combined arms, MC panel, distill seed).
|
||||
- `results-node*/`, `results-node2-final/` — raw syncs from rented H100
|
||||
nodes (500-item eval campaign).
|
||||
- `results-26b/` — **unclear provenance** (Jul 13; layer indices ≤26 mean
|
||||
it is NOT the 26B MoE despite the name — possibly a misnamed early scan).
|
||||
Trust nothing here without re-derivation.
|
||||
- `paper-A/`, `paper-B/`, `paper-D/` — abandoned paper-outline variants
|
||||
(one PLAN.md each); the live outline is PAPER.md itself.
|
||||
- `related_work/` — the two anchor papers (McLeish 2511.07384,
|
||||
Lys 2602.14759), the workspace paper, `relevant_to_us.md` notes,
|
||||
`bibliography.bib`.
|
||||
|
||||
## Conventions worth knowing
|
||||
|
||||
- Every eval prints a `k=0` row first; it must equal the base model
|
||||
bit-exactly (0.488 on MBPP-250) — the sanity anchor that has caught two
|
||||
silent bugs (LESSONS #2, #12).
|
||||
- Difficulty labels (`easy`/`hard`/`drop`) are STaR self-labels:
|
||||
direct-pass / CoT-only-pass / unreachable. "hard" = plan-dependent.
|
||||
- Checkpoints are pre-committed before evals (usually e400); post-hoc
|
||||
checkpoint shopping is flagged as exploratory wherever it happened.
|
||||
- GPU jobs go through the gpuq queue (`gpuq_submit.sh <worker> <job.sh>`),
|
||||
never bare nohup on the Spark; jobs are killed by `pkill -f` self-matches
|
||||
embarrassingly often (LESSONS #6).
|
||||
@@ -62,7 +62,15 @@ reproduction: [`RESULTS.md`](RESULTS.md). Everything on `google/gemma-4-E2B-it`
|
||||
6. **Background jobs must be `setsid`'d** or the harness/session restart
|
||||
kills them mid-run. And `pkill -f <pattern>` will match your own launcher
|
||||
shell if the pattern appears in its command line.
|
||||
7. **Zero-init adapter output layer ⇒ zero grads upstream at step 0** — on
|
||||
7. **Sustained training in a closed cabinet = thermal hard-freezes.** Four
|
||||
crashes in one day (journal stops mid-line, no OOM, no shutdown trace,
|
||||
37GB free at one death) on a DGX Spark that was stable all week under
|
||||
light load. Pattern: only under hours of continuous GPU load; fixed by
|
||||
opening the cabinet. Diagnose by exclusion: earlyoom quiet + journal
|
||||
truncation + load correlation = thermal, not software. And do NOT
|
||||
auto-restart training via @reboot cron on a thermally-suspect box — it
|
||||
risks a crash loop with no human circuit breaker.
|
||||
8. **Zero-init adapter output layer ⇒ zero grads upstream at step 0** — on
|
||||
`mlp[0]` this is expected (LoRA-B-style), not a bug; check the output
|
||||
layer's grad instead.
|
||||
|
||||
@@ -76,3 +84,13 @@ reproduction: [`RESULTS.md`](RESULTS.md). Everything on `google/gemma-4-E2B-it`
|
||||
3. **Verify with the lens, gate with the labels**: the J-lens picks the band,
|
||||
measures whether loops compute, and diagnoses failures; STaR difficulty
|
||||
labels supervise both the curriculum and (next) the adaptive-depth gate.
|
||||
|
||||
## 12. Pin model revisions on fresh nodes
|
||||
Upstream updated google/gemma-4-12B-it mid-project: the new chat template
|
||||
adds a `<|channel>thought` scaffold, and greedy generation closes the empty
|
||||
thought channel and stops — every generation decodes to "". Symptom:
|
||||
0/500 pass rates with rc=0 (looks like a harness bug, is a silent model
|
||||
swap). E2B was unaffected. Fix: `hf download --revision <hash>` + repoint
|
||||
`refs/main` in the cache; node_setup.sh now pins both models (12B
|
||||
0e2b1058…, E2B 9dbdf8a8…). Rule: any cross-node result assumes identical
|
||||
model revisions — pin them, don't trust "main".
|
||||
|
||||
@@ -1,286 +1,519 @@
|
||||
# Retrofitting Latent Planning onto a Frozen Language Model via Workspace Recurrence
|
||||
# Latent Planning by Workspace Recurrence: an Interpretability-Placed Implant, and What It Actually Buys
|
||||
|
||||
*Working draft, 2026-07-14. All experiments: google/gemma-4-E2B-it (frozen), single DGX Spark. Code and artifacts: `~/jspace`.*
|
||||
*Revision draft, 2026-07-14. Base models: google/gemma-4-E2B-it and
|
||||
gemma-4-12B-it, both frozen. Hardware: DGX Spark + rented 2×/8×H100 nodes.
|
||||
Code, per-item logs, and pre-registrations:
|
||||
https://git.draic.info/nils/jspace (public). Statistics:
|
||||
`results-loop/STATS.md`.*
|
||||
|
||||
## Abstract
|
||||
|
||||
Interpretability work with an averaged-Jacobian lens ("J-lens") shows that
|
||||
mid-depth layers of a pretrained language model form a *workspace*: a band of
|
||||
layers that holds verbalizable, unspoken intermediate content. We ask whether
|
||||
that band can be **iterated in place** — spending more serial compute per
|
||||
input without emitting reasoning tokens — on a *frozen* model. A naive loop
|
||||
diverges: the band is not a self-map. We show that a 1.6M-parameter
|
||||
**anchor-dominant merge adapter** (0.03% of the model) at the band entrance
|
||||
makes the recurrence a stable fixed-point iteration, and that training only
|
||||
this adapter — with self-generated, verifier-filtered supervision and a
|
||||
difficulty→depth curriculum — turns iteration into computation. On MBPP,
|
||||
looping the workspace over the prompt ("latent planning") raises pass@1 on
|
||||
plan-dependent problems from **5.5% to 30.9–43.6%** (three seeds, full test
|
||||
set, execution-verified); overall accuracy is unchanged-to-slightly-improved
|
||||
(51.8% → 51.8–53.8%, within noise at n=500) — the method's value is
|
||||
cost-shaped (silent, prefill-parallel, no per-token overhead), not
|
||||
accuracy-dominance. Controls attribute the hard-bucket gain to the
|
||||
recurrence itself: a same-size adapter trained on identical data *without*
|
||||
the loop reaches only 17.9%, exactly matching the untrained loop. On GSM8K the picture inverts — no recurrent variant beats the
|
||||
weights-only control — and a four-arm decomposition localizes why: the loop
|
||||
performs *plan refinement*, which code synthesis needs and answer-time
|
||||
arithmetic does not. The J-lens provides both the intervention's design
|
||||
(where to loop) and its verification (latent concepts sharpen ~8× per
|
||||
converged iteration). Because the looped prompt states are constant during
|
||||
generation, latent planning is prefill-shaped and adds no per-token cost.
|
||||
Interpretability work with an averaged-Jacobian lens ("J-lens") partitions a
|
||||
pretrained language model's depth into regimes, including a mid-depth
|
||||
*workspace* band that holds verbalizable, unspoken intermediate content. We
|
||||
retrofit recurrence onto this band in a **frozen** gemma-4-E2B: a
|
||||
1.6M-parameter anchor-dominant merge adapter (0.03% of parameters) at the
|
||||
band entrance turns the non-self-map band into a stable recurrence, trained
|
||||
with self-generated, verifier-filtered supervision. Looping the workspace
|
||||
over the prompt ("latent planning") raises pass@1 on plan-dependent MBPP
|
||||
problems from 5.5% to 37.5±5.5% over five seeds — pooled across MBPP,
|
||||
HumanEval, and Rust/MultiPL-E, 4.2%→35.6% (McNemar p≈1.5e-10) — with zero
|
||||
visible tokens and zero decode cost. Placement is decisive, not convenient:
|
||||
the gain exists only at the lens-identified boundary (L14), collapsing below
|
||||
it, and KV-sharing structurally nulls entrances above it. Net of the
|
||||
untrained-merge perturbation floor (20.0%), the loop-specific effect
|
||||
survives paired testing (p=0.007).
|
||||
|
||||
## 1. Introduction
|
||||
A two-part attribution program then bounds the mechanism. First, the
|
||||
content is *amortized, not computed*: recurrence-free plan-distillation
|
||||
into the same adapter matches the loop, gains do not stack, and inference
|
||||
depth beyond k≈4 is flat — the state trajectory is an output-stable orbit,
|
||||
not a converging computation (half of prompts' states never converge at
|
||||
cos 0.9995 by k=8, with no difficulty gradient, so no convergence-based
|
||||
early exit falls out). Width rivals depth on code (trained pause registers:
|
||||
36.4%); recurrence is needed where state must evolve (GSM8K carry,
|
||||
Blocksworld planning). Second, a pre-registered regime sweep spanning
|
||||
unconstrained learned recurrence (Huginn-style), spectrally constrained
|
||||
state maps (Parcae-style), per-iteration weights (Bae-style), and learned
|
||||
anchor coefficients shows that **dynamical stability and substrate fidelity
|
||||
are independent dials**: spectral radius governs convergence only (an
|
||||
unconstrained map drifts to ρ≈4.5 with no fit benefit; a constrained one
|
||||
stays at ρ≈0.3 with no fit cost — both lose 17 points of easy-item
|
||||
accuracy), while fidelity is governed by fixed-point *location*, causally
|
||||
isolated to one design choice — tying the input map to the anchor's convex
|
||||
complement, B=(1−α)I. Per-iteration weights strand the gain at trained
|
||||
depths; every regime buys the same hard-bucket gain (36–46%); no regime
|
||||
exceeds the amortization ceiling at this budget. The hand-tuned recipe is
|
||||
thus the measured optimum of its design space, not a lucky point in it.
|
||||
At 12B the anchor coefficient must become state-dependent (3.8K parameters)
|
||||
to preserve the substrate — the one dial that is task- and scale-dependent.
|
||||
Details and exact numbers: §1 and §3.
|
||||
|
||||
Large language models buy reasoning accuracy with emitted tokens: chains of
|
||||
thought give the network more serial passes, at the cost of latency, output
|
||||
tokens, and bandwidth-bound decode. Recurrent-depth architectures (Universal
|
||||
Transformers; DEQs; Huginn, arXiv:2502.05171; Mixture-of-Recursions,
|
||||
arXiv:2507.10524) buy the same serial compute silently — but require
|
||||
(pre)training the recurrence in at scale.
|
||||
|
||||
We investigate a middle path: **retrofit** recurrence onto an off-the-shelf
|
||||
frozen model, using an interpretability signal to decide *where*. The
|
||||
J-lens (from the "verbalizable global workspace" line of work) partitions
|
||||
depth into transduction, sensor, workspace, and motor regimes; the workspace
|
||||
band (L14–30 of 35 in our subject model) holds slowly-varying, unspoken
|
||||
intermediates — e.g. 'spider' before answering "8" to *"the animal that spins
|
||||
webs has how many legs?"*. If the workspace approximates "iterate toward a
|
||||
settled representation", looping it should deepen computation without
|
||||
parameters. The contributions:
|
||||
## 1. What this paper claims
|
||||
|
||||
1. **A minimal retrofit that works**: an anchor-dominant merge
|
||||
(`(1−α)e + α·ŝ + MLP([e;ŝ])`, α=0.3, MLP zero-init, 1.6M params) makes
|
||||
the frozen band a stable, answer-preserving recurrence; training only the
|
||||
merge makes iterations *sharpen* rather than hold.
|
||||
2. **A verified capability gain** on plan-dependent code synthesis, with the
|
||||
full attribution grid (weights / untrained loop / trained loop / pause
|
||||
tokens) showing the recurrence is the active ingredient.
|
||||
3. **A mechanistic boundary**: math inverts the result, and the decomposition
|
||||
(prompt-side vs generation-side × weights vs recurrence) identifies the
|
||||
mechanism as plan refinement, not generic extra compute.
|
||||
4. **Deployment properties**: bit-exact KV-cache-compatible inference (loop
|
||||
once at prefill), a difficulty gate trained free from the labeling
|
||||
pipeline, and economics that improve with model scale.
|
||||
*(One model family, two scales: we state findings as empirical regularities,
|
||||
not laws.)*
|
||||
|
||||
1. **A placement regularity.** The retrofit works if and only if the
|
||||
recurrence enters at the lens boundary. Entrances at L9–L13 (same
|
||||
adapter, data, curriculum) destroy overall accuracy (14–34% vs 52%)
|
||||
while recovering at most half the hard-bucket gain; entrance at L14
|
||||
preserves overall and maximizes the gain (fig_placement). Entrances at
|
||||
L17/L24 are *structurally null* in this architecture: KV-sharing makes
|
||||
layers ≥15 reuse keys/values computed at ≤14, so k>0 is bit-identical to
|
||||
k=0 — a hazard for any retrofit method that skips the mechanistic check.
|
||||
Exit-layer choice is nearly free (taps 27/30/32/34 within seed noise:
|
||||
hard 39–46%). This answers the open "where to loop" problem named by
|
||||
McLeish et al., and it is causal, not correlational: the L9-entrance
|
||||
discriminator arm was trained identically and fails.
|
||||
|
||||
2. **A verified, statistically solid capability gain on a narrow slice.**
|
||||
Plan-dependent items (the model solves them with an explicit written plan
|
||||
but not directly): seed-mean 37.5±5.5 on MBPP (best 43.6%); pooled across
|
||||
three benchmarks, 4.2%→35.6%, p≈1.5e-10. Overall accuracy is
|
||||
statistically unchanged on MBPP (p=0.34) and improved on HumanEval
|
||||
transfer (58.5%→66.5%, p=0.011). Net of the untrained-merge floor
|
||||
(20.0% at n=55), the loop-specific effect is +17.5 points (seed mean)
|
||||
and **survives the paired test** (loop vs untrained merge on hard,
|
||||
p=0.007; distill vs untrained, p=0.0075).
|
||||
|
||||
3. **A deflationary mechanism finding.** The trained loop converges to a
|
||||
fixed point by k≈3–4 and behaves as *amortized plan content*, not
|
||||
iterative computation: plan-distillation into the identical architecture
|
||||
without recurrence matches it; stacking buys nothing (loop-training a
|
||||
distill-warmed adapter: 34.5%, below distill alone; running the distilled
|
||||
adapter in loop mode: drops to 20.0%); deeper k at inference is flat
|
||||
(k=8: 40.0%; output-stable despite residual state drift, §3.8). The
|
||||
recurrence is a *training-time scaffold* that lets the
|
||||
adapter find plan-shaped content — content that can equally be put there
|
||||
by distillation if plans are available.
|
||||
|
||||
4. **A width-vs-depth pattern.** Trained pause registers (width) capture most of
|
||||
the plan effect on code; recurrence (depth) is needed only where a state
|
||||
must *evolve* — on GSM8K generation-side carry beats registers, and on
|
||||
Blocksworld (pure planning, no world knowledge) the loop lifts hard-split
|
||||
plans 0%→43% at 2B where everything else fails. Plans are wide; execution
|
||||
is deep.
|
||||
|
||||
5. **Honest economics.** The implant's costs: ≈2.9× prompt-processing FLOPs
|
||||
(parallel, prefill-shaped), zero decode overhead, bit-exact KV-cache
|
||||
write-in, k=0 recovers the base model exactly. Its competition at
|
||||
comparable compute (accounting in Appendix A — FLOPs, wall-clock, and
|
||||
token budget do not rank the arms the same way): *oracle* best-of-3
|
||||
(verifier-assisted) wins overall accuracy (57.8%, paired p=0.020 vs
|
||||
loop), but the *deployable* logprob-selected variant drops to 55.0%
|
||||
overall / 27.3% hard — indistinguishable from the latent arms overall
|
||||
and directionally behind on hard. A 50-token visible plan ties the hard
|
||||
bucket. The value proposition without a verifier: no visible tokens, no
|
||||
decode latency, and the hard-slice specialization (distill's 46% >
|
||||
budget-CoT's 40% > deployable sampling's 27%).
|
||||
|
||||
6. **Scale transfers only with a state-dependent stability dial.** At 12B the
|
||||
2B-tuned constant α=0.3 collapses overall accuracy (72.6%→43.0%); the
|
||||
damage is present *before* adapter training (untrained-loop arm) and is
|
||||
not fixed by retuning α or LR. A per-position learned coefficient
|
||||
α=σ(w·[e;ŝ]+b) restores MBPP (overall 69.4%, hard 11.4%→27.3%) — but
|
||||
fails to rescue Blocksworld-12B and yields only a nominally positive,
|
||||
not significant GSM8K-12B overall delta (35.9%→36.7% at k=1, p=0.86).
|
||||
|
||||
## 2. Method
|
||||
|
||||
**Locating the band.** The lens reads residual state h at layer ℓ through the
|
||||
averaged Jacobian J̄_ℓ = E[∂h_final/∂h_ℓ] and the unembedding. Depth regimes
|
||||
follow from what the readout tracks (input echo / abstract content / output
|
||||
token). On gemma-4-E2B: workspace ≈ L14–30 (1.08B params, 58% of decoder).
|
||||
averaged Jacobian J̄_ℓ = E[∂h_final/∂h_ℓ] and the unembedding; depth regimes
|
||||
follow from what the readout tracks. On gemma-4-E2B: workspace ≈ L14–30 of
|
||||
35; on 12B: L36–45 of 48.
|
||||
|
||||
**Making the band a self-map.** Feeding L30's output to L14 collapses in one
|
||||
step (out-space ≠ in-space; norms and content 17 layers "downstream").
|
||||
Additive anchoring diverges. The fix is DEQ-style input injection done by
|
||||
hand: with e = L13's output (fixed anchor) and s the fed-back band output,
|
||||
**Making the band a self-map.** Feeding L30's output back to L14 collapses
|
||||
(out-space ≠ in-space). With e = L13's output (fixed anchor) and s the
|
||||
fed-back, norm-matched band output:
|
||||
|
||||
L14-in = (1−α)·e + α·(s · |e|/|s|) + MLP([e ; s·|e|/|s|]), α = 0.3.
|
||||
L14-in = (1−α)·e + α·ŝ + MLP([e ; ŝ]), ŝ = s·‖e‖₂/‖s‖₂
|
||||
|
||||
Zero-initializing the MLP's output layer makes the untrained adapter exactly
|
||||
the hand merge, which is stable and answer-preserving for ≥11 iterations but
|
||||
only *holds* content (lens concept flat).
|
||||
(per-position L2 norms over the hidden dimension, computed in fp32)
|
||||
|
||||
**Training only the merge.** Supervision is self-generated and
|
||||
verifier-filtered (STaR-style): the frozen model attempts each task directly
|
||||
and with explicit planning/CoT; items it solves only with planning are
|
||||
"hard", direct solves "easy", neither "drop". Cross-entropy on answer/code
|
||||
tokens of the *direct* prompt; the model's own verified outputs are the
|
||||
targets (in-distribution). A **difficulty→depth curriculum** trains easy
|
||||
items at loop depth k=1, mixed at k=2, hard only at k=2–4, so loss on hard
|
||||
items is reducible only through the recurrence. For generation tasks the
|
||||
loop applies to the **prompt span only** ("latent planning"): generated
|
||||
tokens run the plain path but attend to the looped prompt states; this
|
||||
removes exposure bias structurally.
|
||||
α=0.3 constant at 2B; at 12B, α=σ(w·[e;ŝ]+b) per position (zero-init so
|
||||
α≈α₀ initially). MLP output zero-init: the untrained adapter is exactly the
|
||||
hand merge — stable, answer-preserving, content-holding.
|
||||
|
||||
**Inference cost.** Causality makes the looped prompt states independent of
|
||||
generated tokens, so they are computed once; a hooked prefill writes them
|
||||
into the KV cache and generation proceeds natively (verified bit-identical;
|
||||
≥3.5× faster than recomputation). The concrete overhead at k=4 is 5 passes
|
||||
over the band's 17/35 layers at prefill — ≈2.9× prompt-processing FLOPs,
|
||||
parallel across positions — and **zero** additional decode cost. Explicit
|
||||
planning with ~200 emitted tokens costs more total FLOPs and pays them
|
||||
serially at bandwidth-bound decode; this asymmetry grows with model size.
|
||||
**Training.** STaR-style self-labeling: items the frozen model solves only
|
||||
with an explicit plan/CoT are "hard", direct solves "easy", neither "drop".
|
||||
Cross-entropy on answer/code tokens of the direct prompt, model's own
|
||||
verified outputs as targets. Difficulty→depth curriculum (easy k=1, mixed
|
||||
k=2, hard k=2–4). The loop applies to the **prompt span only**; generated
|
||||
tokens run the plain path but attend to looped prompt states. Variants
|
||||
trained the same way: **pause-N** (N trained register tokens appended to the
|
||||
prompt, no recurrence), **plan-distill** (KL from the model's own
|
||||
plan-in-context distribution into the FF adapter), **rung-2** (warm-started
|
||||
adapter + entrance-faded LoRA rank 8 on the band's first layers, loop-only
|
||||
via a global toggle), and **stack** arms (distill-warm + loop training;
|
||||
distilled adapter evaluated in loop mode).
|
||||
|
||||
**Truncated backprop is certified by contraction.** Recurrent-regime arms
|
||||
train with tail-only BPTT (gradients through the last 4 iterations; the
|
||||
no-grad prefix stores no activations, so memory is constant in depth).
|
||||
The truncation bias scales as ρ(A)^(T−tail) — at ρ=0.3 the discarded terms
|
||||
are ≤1%, making the cheap estimator essentially exact; at ρ≥1 it is
|
||||
dominated by what it discards. Stability, fixed-point convergence, valid
|
||||
tail gradients, and the convergence-halting exit signal are all the same
|
||||
dial.
|
||||
|
||||
**Inference.** Looped prompt states are causally independent of generated
|
||||
tokens: computed once at prefill, written into the KV cache by a hooked
|
||||
forward pass, generation native. Verified bit-identical to the slow path.
|
||||
Cost at k=4: ≈2.9× prefill FLOPs, **zero** decode overhead.
|
||||
|
||||
## 3. Results
|
||||
|
||||
### 3.1 Latent planning on code (MBPP)
|
||||
Statistics throughout: Wilson 95% CIs; paired comparisons by exact McNemar;
|
||||
all headline arms evaluated on the full 500-item MBPP test split (hard
|
||||
bucket n=55), HumanEval n=164 (hard n=38), Rust/MultiPL-E n=154 (hard n=25),
|
||||
execution-verified.
|
||||
|
||||
Full 500-item test split, greedy decode, unit-test-verified. Hard bucket =
|
||||
items the frozen model solves only with an explicit written plan (n=55).
|
||||
**Bucket definition, stated up front.** Hard labels come from labeling runs
|
||||
of the frozen base model on the test items themselves (direct vs
|
||||
plan-in-context, greedy). This is legitimate for *descriptive* slicing but
|
||||
would be circular for selection — so no arm, hyperparameter, checkpoint, or
|
||||
loop depth was ever chosen using bucket results (pre-registered;
|
||||
`PROTOCOL_UNIFIED.md` items 1–2, 8). Because the bucket conditions on k=0
|
||||
failure, regression-to-mean inflates *any* intervention's bucket score: the
|
||||
untrained merge already reaches ~18%, and we therefore report the
|
||||
loop-specific effect **net of that floor** wherever attribution is claimed.
|
||||
Robustness: redefining "hard" as labeled-hard ∧ k=0-fails-in-all-five-seeds
|
||||
(52/55 items) moves headline numbers <2 points; both definitions share the
|
||||
base model, which an independent difficulty proxy would not — we flag this
|
||||
as an open external check.
|
||||
|
||||
| k=4 (prompt-only loops) | hard pass@1 | overall |
|
||||
### 3.1 The placement regularity
|
||||
|
||||

|
||||
|
||||
Entrance-layer sweep with everything else fixed. L14 (lens boundary):
|
||||
hard 43.6%, overall 53.6%. L13: hard 17.9%, overall 34.4%. L9–L12: overall
|
||||
14.0–30.8% (substrate destroyed). L17/L24 entrances: k>0 ≡ k=0 (KV sharing;
|
||||
verified bit-identical) — the 12B model has no shared-KV layers, making it
|
||||
the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34):
|
||||
hard 39.3–46.4%, within seed spread. The lens boundary is necessary; the
|
||||
exit is a free parameter — the completed five-point exit sweep
|
||||
(L23/27/30/32/34) spans hard 39.3–46.4% with L23 at the top (46.4% at k=2),
|
||||
all within seed spread.
|
||||
|
||||
### 3.2 The attribution ladder
|
||||
|
||||

|
||||
|
||||
MBPP hard bucket (plan-dependent, n=55 unless noted):
|
||||
|
||||
| arm | hard pass@1 | overall |
|
||||
|---|---|---|
|
||||
| baseline (k=0) | 5.5% | 51.8% |
|
||||
| trained loop, seed 0 | **43.6%** | 53.6% |
|
||||
| trained loop, seed 1 | **41.8%** | 53.8% |
|
||||
| trained loop, seed 2 | **30.9%** | 51.8% |
|
||||
| base (k=0, bit-exact) | 5.5% | 51.8% |
|
||||
| untrained loop (α-merge only, k=4) | 20.0% | 50.2% |
|
||||
| trained FF, no recurrence (k=1) | 27.3% | 53.6% |
|
||||
| pause-16 registers (width) | 36.4% | 55.2% |
|
||||
| **trained loop k=4** (seed mean, 5 seeds) | **37.5±5.5** (best 43.6) | 53.6% |
|
||||
| rung-2: + entrance-faded band LoRA (n=28) | 42.9/46.4 (2 seeds) | 51.2/52.4 |
|
||||
| **plan-distilled FF** (mean, 8 runs) | **45.7±4.6** (best 49.1) | 55.5% |
|
||||
| budget-CoT (50 visible tokens) | 40.0% | 53.8% |
|
||||
| best-of-3 sampling (≈matched FLOPs) | 32.7% | **57.2%** |
|
||||
| explicit plan in context (ceiling) | 94.5% | 59.0% |
|
||||
|
||||
Silent loops recover roughly 40% of what explicit planning achieves, at zero
|
||||
visible-token cost, with no overall regression (the easy-item perturbation
|
||||
tax, ~9 points, is offset by hard/drop gains; a gate removes most of it,
|
||||
§3.4).
|
||||
Significance structure (McNemar, `STATS.md`): loop vs base on hard,
|
||||
p=5.7e-6; every latent-arm-vs-latent-arm difference (loop vs distill, distill
|
||||
vs stack) is **not significant** at n=55; loop vs base *overall* is not
|
||||
significant on MBPP (p=0.34). The ladder's shape is reliable; its fine
|
||||
ordering is not.
|
||||
|
||||

|
||||
**Net accounting.** The attribution-critical comparison is trained-loop vs
|
||||
*untrained merge*, not vs base: gross 5.5→37.5 (seed mean), of which the
|
||||
untrained perturbation floor is 20.0 points — the loop-specific net is
|
||||
+17.5 (seed mean) / +23.6 (best seed), and the paired item-level test is
|
||||
decisive (loop-only 17, untrained-only 4, p=0.007; distill likewise
|
||||
p=0.0075). The trained-FF control (27.3%) sits between floor and loop,
|
||||
not significantly above the floor (p=0.48): weights alone buy little
|
||||
without either recurrence or plan supervision. All controls now n=500 /
|
||||
hard n=55, same harness.
|
||||
|
||||
### 3.2 Attribution: the recurrence is the ingredient
|
||||
### 3.3 The decisive tests: nothing stacks
|
||||
|
||||
250-item subset; same data, same 1.6M parameters, same insertion point:
|
||||
If the loop performed genuine iterative computation, plan-distilled content
|
||||
plus recurrence should compound. It does not:
|
||||
|
||||
| arm | hard pass@1 |
|
||||
|---|---|
|
||||
| baseline | 3.6% |
|
||||
| untrained loop (α-merge only) | 17.9% |
|
||||
| trained adapter, **no loop** (weights control) | 17.9% |
|
||||
| trained **loop** | **42.9–46.4%** |
|
||||
- **Distill-warm + loop training**: hard 34.5% — below distill alone.
|
||||
- **Distilled adapter run in loop mode**: hard 20.0%, overall 45.8% —
|
||||
looping *degrades* the distilled weights.
|
||||
- **Pause-16 + distill**: hard 30.9% — no width stacking either.
|
||||
- **Inference depth beyond convergence**: k=8 hard 40.0% ≈ k=4 (fixed point,
|
||||
cos(sₖ,sₖ₋₁)=1.000 by k≈3–4).
|
||||
|
||||
The weights control lands exactly on the untrained-loop value: ~18 points is
|
||||
what perturbation-plus-format-alignment buys. The remaining ~28 points
|
||||
require iterating the band. Post-hoc depth selection is excluded by
|
||||
pre-registration (k=2 fixed on validation before test numbers existed;
|
||||
k-curves reported descriptively).
|
||||
Reading: the recurrence is a **training-time scaffold**. The curriculum
|
||||
forces hard-item loss to be reducible only through the loop, and what the
|
||||
adapter learns to inject is plan-shaped content — the same content
|
||||
distillation installs directly when explicit plans are available. The loop's
|
||||
distinctive value is that it finds this content *without* plan supervision
|
||||
(STaR labels only say which items needed plans, not what the plans were).
|
||||
|
||||
**Checkpoint selection.** No checkpoint was chosen using test or generation
|
||||
results. Seed 0's checkpoint (step 399) was fixed at training time from the
|
||||
validation-CE overfitting inflection, before any generation eval of that
|
||||
adapter; seeds 1–5 use step 400 by pre-commitment made before those seeds
|
||||
were trained. We separately report that validation CE is a poor proxy for
|
||||
generation accuracy (a checkpoint selected by val-CE on a sibling arm
|
||||
underperformed a later one), which is why the fixed-step rule is used
|
||||
rather than per-seed val selection.
|
||||
### 3.4 Compute-matched honesty
|
||||
|
||||
### 3.3 The boundary: math
|
||||
At approximately matched FLOPs (Appendix A gives the accounting, separated
|
||||
into FLOPs, wall-clock, and token budget), the token-space comparison
|
||||
splits into two very different claims:
|
||||
|
||||
On GSM8K, *no* recurrent variant beats the weights-only control. The four-arm
|
||||
grid (hard bucket) decomposes the failure:
|
||||
| best-of-3 variant | overall | hard | vs loop (paired) |
|
||||
|---|---|---|---|
|
||||
| **oracle** (any-of-3 passes; needs a perfect verifier) | 57.8% | 34.5% | beats loop overall, p=0.020 |
|
||||
| **deployable** (highest mean logprob of 3) | 55.0% | 27.3% | n.s. overall (p=0.47); loop ahead on hard 16–7 (p=0.09) |
|
||||
|
||||
| GSM8K hard | prompt-side only | touches generation |
|
||||
|---|---|---|
|
||||
| feedforward weights | **11.8%** | 4.7% (pause-token control) |
|
||||
| recurrence | 6.3–8.7% (prompt loop) | 9.4% (cross-token carry) |
|
||||
The earlier draft's "sampling wins overall" was the **oracle** number — an
|
||||
upper bound requiring an external verifier that MBPP's own tests provide
|
||||
but a deployment does not. With the deployable selector (identical seeded
|
||||
samples, so the comparison is exact), best-of-3 is statistically
|
||||
indistinguishable from the latent arms overall, *behind* them
|
||||
directionally on the hard bucket, and pays ≈3× visible tokens and serial
|
||||
decode for it. Budget-CoT-50 remains the strongest honest token baseline
|
||||
(53.8% overall, hard 40.0%; per-item rerun 53.8/38.2) — and the paired
|
||||
tests confirm it is a *tie* with the latent arms on both axes (p≥0.69 vs
|
||||
loop and distill), at the price of 50 visible tokens and their serial
|
||||
decode latency. The implant's advantages at matched compute
|
||||
without a verifier: zero visible tokens, zero decode overhead, and the
|
||||
hard-slice crown under distillation (46%). Where a task *does* come with a
|
||||
cheap verifier, oracle-style sampling is the better overall-accuracy
|
||||
spend — both halves belong in the deployment picture.
|
||||
|
||||
Orthogonal effects: perturbing free-running generation positions is costly
|
||||
for either mechanism; recurrence beats weights only where a state must
|
||||
evolve (the generation side — carry doubles the pause control in-harness),
|
||||
and loses on the static prompt side. No variant beats the 10.5% overall
|
||||
baseline. Reading: the trained loop performs **plan refinement**; code
|
||||
synthesis is plan-shaped, multi-step arithmetic is not — its serial
|
||||
computation happens during the answer, and one frozen band pass per token
|
||||
cannot perform it silently at 2B. CoT tokens remain load-bearing for math.
|
||||
(Hard-bucket cells carry an outcome-selection caveat — buckets were defined
|
||||
by greedy baseline outcomes; sampled relabeling is in progress — so the math
|
||||
conclusion is stated on overall numbers.)
|
||||
### 3.5 Width vs depth, and the task boundary
|
||||
|
||||
### 3.4 Mechanism and deployment
|
||||
Pause registers (width) reach 36.4% (16 registers; 8: 30.9%, 32: 34.5% — flat
|
||||
in N) on MBPP hard: static plan content fits in registers. GSM8K inverts the
|
||||
prompt-side result entirely (no variant beats the weights control
|
||||
prompt-side), but generation-side *carry* — recurrence across token steps —
|
||||
doubles the pause control on hard items: arithmetic's serial state evolves
|
||||
during the answer. Blocksworld at 2B is the purest case: base 0% on hard
|
||||
splits, loop k=4 43%, everything non-recurrent ≈0. The pattern: **plans are
|
||||
wide; execution is deep.** Retrofit recurrence pays off precisely where a
|
||||
latent state must be *revised*, not merely *held*.
|
||||
|
||||
**Fixed point.** The trained loop takes a large first step
|
||||
(cos(s₁,s₀)=0.926 vs 0.977 untrained) and converges bit-exactly by k≈3–4
|
||||
(cos=1.000), where accuracy and lens-sharpening plateau — extra iterations
|
||||
are no-ops, explaining the k-curve shape.
|
||||
### 3.6 Scale: the stability dial
|
||||
|
||||

|
||||

|
||||
|
||||
**Lens verification.** P(latent concept) under the J-lens at the band exit
|
||||
rises 0.015→0.13 across iterations after training (~8× the untrained
|
||||
control, which only holds). The same lens that located the band verifies
|
||||
that looping deepens its computation — and makes the silent reasoning
|
||||
inspectable.
|
||||
At 12B (no shared KV — unconfounded), constant α=0.3: overall collapses
|
||||
72.6%→43.0% at k=4 while hard limps to 11.4%. The untrained-loop arm shows
|
||||
the damage precedes adapter training; α=0.15 and LR retuning do not fix it
|
||||
(47.6/52.6% overall). The state-dependent coefficient does, on MBPP:
|
||||
overall 69.4% (base 72.4%), hard 11.4%→27.3%. It does **not** rescue
|
||||
Blocksworld-12B (easy items destroyed at k=4; constant-α had reached hard
|
||||
40% but also destroyed easy) and yields a **nominally positive, not
|
||||
significant** overall delta on GSM8K-12B (35.9→36.7 at k=1; paired McNemar
|
||||
on 32 discordant items, p=0.86; hard 1.6→10.6) — no arm anywhere in the
|
||||
program produced a statistically significant overall gain at 12B.
|
||||
Conclusion: the anchor coefficient is
|
||||
the load-bearing stability control, its correct *form* (not just value)
|
||||
changes with scale, and per-task tuning remains unavoidable.
|
||||
|
||||
**Gate.** A logistic probe on the k=0 workspace state (supervised for free
|
||||
by the STaR labels) routes prompts: predicted-easy at k=0, predicted-hard at
|
||||
k=4. Result: overall equal to the best uniform depth with easy items fully
|
||||
preserved (97.5% vs 98.4% baseline); probe precision (19% at 64% recall) is
|
||||
the current ceiling.
|
||||
### 3.7 Transfer: substrate, not task
|
||||
|
||||
**Negative results with content.** Mixed-task (code+math) training regressed
|
||||
both tasks versus dedicated adapters, despite indistinguishable validation
|
||||
CE — cross-entropy parity does not predict generation parity. Validation-CE
|
||||
checkpoint selection likewise failed to track generation accuracy.
|
||||

|
||||
|
||||
MBPP-trained implants applied unchanged: **HumanEval** overall 58.5%→66.5%
|
||||
(loop k=4, p=0.011 vs base; hard 0→31.6%). The decisive control: the
|
||||
*untrained* merge already reaches 64.6%, and trained-vs-untrained is **not
|
||||
significant** (paired McNemar at k=2, 9 vs 7 discordant, p=0.80). What
|
||||
transfers significantly is the *merge perturbation itself*, not the
|
||||
MBPP-trained content — the cleanest evidence that off-distribution value is
|
||||
substrate-shaped rather than task-memorized. (The transferred pause adapter
|
||||
reaches 66.5%, hard 38.9%, consistent with the same reading.)
|
||||
|
||||
**LiveCodeBench sharpens this into a dissociation** (150 newest stdin
|
||||
problems, Nov 2024–Apr 2025, execution-verified; no LCB training anywhere
|
||||
in the pipeline; base 18.7%):
|
||||
|
||||
| arm (MBPP-trained where trained) | overall | hard (n=25) | vs base, paired |
|
||||
|---|---|---|---|
|
||||
| **untrained merge, k=4** | **24.0%** | **36.0%** | **+**, p=0.039 |
|
||||
| trained loop, k=4 | 15.3% | 8.0% | −, p=0.23 |
|
||||
| distill FF, k=1 | 12.7% | 16.0% | **−**, p=0.049 |
|
||||
|
||||
Far from distribution, the *trained content is a liability* (distill
|
||||
significantly hurts; untrained-vs-trained-loop is 14–1 discordant,
|
||||
p=0.001) while the *untrained anchored recurrence significantly helps* —
|
||||
the training-free regime of Lys et al. is the right choice off-distribution,
|
||||
and the amortized-content reading of §3.3 predicts exactly this: what the
|
||||
adapter learned is MBPP-shaped plan content, valuable where plans look like
|
||||
MBPP plans and harmful where they don't. Transfer ordering by distance:
|
||||
HumanEval (near) — trained ≈ untrained; Rust (mid) — trained helps the hard
|
||||
bucket; LCB (far) — untrained wins outright. Caveats: single seed per arm,
|
||||
hard n=25, one benchmark at the far end. **Rust/MultiPL-E** (Python-trained, different
|
||||
language, compile-run-verified): hard 8.0%→24.0% (p=0.125 at n=25 —
|
||||
directionally consistent, underpowered). **Blocksworld** MBPP-transfer:
|
||||
hard 0→14.3% (task-trained: 43%). Content transfers where the substrate's
|
||||
plan-representation overlaps; task-specific training still dominates.
|
||||
|
||||
### 3.8 Mechanism, verification, deployment
|
||||
|
||||
The trained loop takes a large first step (cos(s₁,s₀)=0.926 vs 0.977
|
||||
untrained); accuracy and lens-sharpening plateau by k≈3–4. A population
|
||||
probe (n=250, state-cosine threshold 0.9995) shows the plateau is
|
||||
*output-level*: half the prompts' states are still drifting at 1e-3–1e-4
|
||||
cosine scale at k=8 while generation is already depth-stable — an
|
||||
output-stable orbit rather than a literal state fixed point, with no
|
||||
difficulty gradient in state-convergence depth. Consequently,
|
||||
convergence-based early exit ("free ACT") does not fall out of the state
|
||||
trajectory; halting would need an output-level signal. P(latent concept) under the J-lens at the band exit rises
|
||||
0.015→0.13 across iterations (~8× the untrained hold) — the lens that placed
|
||||
the implant also renders its silent content inspectable. The STaR labels
|
||||
train a free difficulty gate (route predicted-hard to k=4, else k=0);
|
||||
gate quality (19% precision at 64% recall) is the current ceiling on
|
||||
removing the easy-item perturbation tax. k=0 is the exact base model by
|
||||
construction — the implant is removable at token granularity.
|
||||
|
||||
**General-capability panel** (ARC-Challenge, WinoGrande, HellaSwag, MMLU;
|
||||
800 items each, length-normalized MC likelihood via the chat template, loop
|
||||
applied to the context span). The safety answer is clean — **k>0 does not
|
||||
damage general abilities**:
|
||||
|
||||
| arm | ARC-C | WinoGrande | HellaSwag | MMLU |
|
||||
|---|---|---|---|---|
|
||||
| base (k=0) | 36.0 | 55.9 | 52.3 | 30.1 |
|
||||
| loop k=2 (MBPP adapter) | 36.1 | 55.3 | 49.6 | 31.3 |
|
||||
| distill FF (MBPP) | 41.8 | 56.6 | 57.0 | 31.8 |
|
||||
|
||||
The loop arm is flat within noise (largest move −2.6 on HellaSwag,
|
||||
unpaired n=800). The distill adapter *nominally improves* every benchmark
|
||||
(+5.8 ARC, +4.8 HellaSwag) — consistent with §3.7's finding that these
|
||||
implants carry a generically useful perturbation component, though
|
||||
MC-likelihood scoring and generation quality are different regimes (see
|
||||
the LCB result below before reading this as free capability).
|
||||
|
||||
### 3.9 Negative results with content
|
||||
|
||||
Mixed-task (code+math) training regressed both tasks at equal validation CE
|
||||
— CE parity does not predict generation parity, and validation-CE checkpoint
|
||||
selection fails likewise (fixed-step pre-commitment used instead; no
|
||||
checkpoint was selected on test or generation results). GSM8K distillation
|
||||
collapsed to empty outputs twice (E2B first attempt, 12B) on 3-token targets
|
||||
under KL-dominant loss; a CE-dominant retry at E2B trained but reached only
|
||||
hard 4.7%. Plan-distillation on GSM8K underperforms its MBPP twin even when
|
||||
training succeeds: consistent with §3.5, there is little static plan content
|
||||
for math to amortize.
|
||||
|
||||
## 4. Related work
|
||||
|
||||
Two recent papers bracket this work. **McLeish et al. (arXiv:2511.07384)**
|
||||
retrofit depth-recurrence into pretrained 1B models via layer surgery +
|
||||
continued pretraining (~50B tokens, all parameters, Muon, recurrence
|
||||
curriculum to r=32): the generic claims "retrofitted recurrence works and
|
||||
beats the non-recurrent parent" and "pretrain-then-convert" are theirs, at
|
||||
~5 orders of magnitude more training cost than ours. They name layer choice
|
||||
as an open problem; our lens-derived band with its causal backing (anchor
|
||||
cliff at L14, tap invariance, wrong-band ≈ 0, KV-sharing hazard) is a direct
|
||||
answer to it. Unlike their surgery (which needs a healing phase), our k=0
|
||||
exactly recovers the base model. **Lys et al. (arXiv:2602.14759)** loop
|
||||
frozen models training-free and show naive looping degrades (distribution
|
||||
shift) while interpolating with the un-looped state rescues it — independent
|
||||
convergent evidence for our anchor-dominant merge; their whole setting
|
||||
corresponds to the untrained cell of our attribution table (17.9% hard =
|
||||
our FF/untrained level), evaluated by likelihood rather than execution.
|
||||
**McLeish et al. (arXiv:2511.07384)** retrofit depth-recurrence via layer
|
||||
surgery + ~50B-token continued pretraining of all parameters; they name
|
||||
layer choice as an open problem — §3.1 is a causal answer. Their surgery
|
||||
needs a healing phase; our k=0 is exactly the base model. **Lys et al.
|
||||
(arXiv:2602.14759)** loop frozen models training-free; their finding that
|
||||
naive looping degrades while interpolation with the un-looped state rescues
|
||||
it is independent convergent evidence for anchor-dominance, and their
|
||||
setting is the untrained cell of our ladder (17.9%).
|
||||
|
||||
**One mechanism, three regimes.** All three works are variants of a single
|
||||
design: mix the fed-back state with an anchor derived from the un-looped
|
||||
computation. Lys et al.'s inference-time moving average η·h⁽⁰⁾+(1−η)·h⁽ᵗ⁾ is
|
||||
an untrained anchor-mixing coefficient; our (1−α)e + α·ŝ + MLP([e;ŝ]) is its
|
||||
trained analogue (fixed mix + learned correction); McLeish et al.'s
|
||||
concatenated input injection is the fully learned limit, trained end-to-end.
|
||||
The anchor coefficient is the stability dial of frozen-band looping: Lys
|
||||
et al.'s naive-looping collapse is the zero-anchor (α→1) limit, their
|
||||
regularization gains are the untrained anchored regime, and our 12B failure
|
||||
at the 2B-tuned α=0.3 — with the untrained-substrate arm showing the damage
|
||||
is pre-training-of-the-adapter — is the same dial mis-set at a new scale.
|
||||
Stability of retrofitted recurrence appears to be governed by how strongly
|
||||
the loop is anchored, across all three training budgets.
|
||||
**One mechanism, three regimes.** All three works mix the fed-back state
|
||||
with an anchor from the un-looped computation. Lys et al.'s moving average
|
||||
η·h⁽⁰⁾+(1−η)·h⁽ᵗ⁾ is an untrained anchor coefficient; our
|
||||
(1−α)e + α·ŝ + MLP is its trained analogue; McLeish et al.'s input injection
|
||||
is the fully-learned limit. The 12B episode closes the loop on this
|
||||
unification: the coefficient is the stability dial, naive looping is its
|
||||
α→1 collapse limit, and our scale failure + state-dependent fix show the
|
||||
dial must itself become a function of the state as models grow. Our stacking
|
||||
results add a caution for the whole family: if retrofitted recurrence
|
||||
content is amortizable (§3.3), some of the family's gains may be
|
||||
reproducible by distillation without inference-time recurrence — a control
|
||||
neither bracket paper runs.
|
||||
|
||||
Earlier lineage: Universal Transformers (adaptive depth); DEQ (fixed-point
|
||||
inference); Huginn (arXiv:2502.05171) — prelude/core/coda from scratch;
|
||||
Mixture-of-Recursions (arXiv:2507.10524) — learned per-token depth; Relaxed
|
||||
Recursive Transformers (arXiv:2410.20672) — uptrained tied layers; Coconut —
|
||||
latent CoT; pause tokens (Goyal et al.) — token-space silent compute, whose
|
||||
trained-adapter variant proved a near-match for our loop on MBPP (§3.2).
|
||||
Earlier lineage: Universal Transformers; DEQ; Huginn (2502.05171);
|
||||
Mixture-of-Recursions (2507.10524); Relaxed Recursive Transformers
|
||||
(2410.20672); Coconut; pause tokens (Goyal et al.) — whose trained variant
|
||||
proved a genuine rival, not a strawman (§3.2, §3.5).
|
||||
|
||||
What remains distinct here: **interpretability-derived loop placement with
|
||||
causal validation** (answering McLeish et al.'s open problem); **a 1.6M-param
|
||||
trained merge on a fully frozen base** (between Lys et al.'s free end and
|
||||
McLeish et al.'s full-retraining end, and the only one of the three where
|
||||
the base model is provably untouched); **prompt-only latent planning with
|
||||
bit-exact KV-cache write-in and zero decode cost**; **the attribution
|
||||
ladder** (untrained / weights / pause / loop / explicit plan) — neither
|
||||
bracket paper runs compute-matched token-space controls; and **difficulty-
|
||||
adaptive depth via the STaR-label gate**, named as future work in both.
|
||||
**Saunshi et al. (2025)** argue looped transformers trade composition
|
||||
against memorization: looping buys iterative reasoning, not fact storage.
|
||||
Our results reproduce this axis *within one frozen model*: k>0 moves only
|
||||
the plan-dependent (compositional) slice, leaves recall-flavored MC
|
||||
benchmarks flat (§3.8), and the content-injecting distill arm — not the
|
||||
loop — is what nudges knowledge benchmarks up. Their looping-based
|
||||
regularization (loop harder on reasoning, relax for retrieval) has an
|
||||
inference-time analogue in our difficulty gate: route predicted
|
||||
plan-dependent prompts to k=4 and everything else to k=0, which is the
|
||||
exact base model. Retrofit looping makes the composition/memorization
|
||||
trade a *per-prompt routing decision* instead of a pretraining commitment.
|
||||
|
||||
What remains distinct here: interpretability-derived placement with causal
|
||||
validation; a fully frozen base with bit-exact k=0 and zero-decode-cost KV
|
||||
write-in; the complete attribution ladder including compute-matched
|
||||
token-space baselines and stacking tests; the width/depth task pattern; and the
|
||||
amortizability finding itself.
|
||||
|
||||
## 5. Limitations
|
||||
|
||||
One base model family at 2B-effective scale (12B replication in progress);
|
||||
two task families. **Location specificity is not yet ablated**: a
|
||||
pre-registered control looping shifted/early/late/width-matched bands with
|
||||
identical adapter and curriculum is queued; until it lands, the results are
|
||||
formally consistent with "any wide mid-depth band works", and the lens claim
|
||||
rests on discovery convenience plus mechanism verification. Hard buckets are
|
||||
small (n=55 greedy / n=33 sampled) with seed spread of ±6 items; sampled
|
||||
relabeling shows 97% agreement with greedy labels, and intervals accompany
|
||||
all bucket cells in the final tables. The MBPP attribution grid lacks a
|
||||
pause-token arm and a plan-distillation baseline (both queued) — the GSM8K
|
||||
grid has the former. Easy-item perturbation tax is not eliminated (gate
|
||||
preserves easy items but probe precision is 19%). Visible planning remains
|
||||
stronger on absolute accuracy — the claim is cost-and-latency-shaped.
|
||||
**Mixed-task training regressed both tasks**, so the current recipe yields
|
||||
per-task adapters, not one general silent-planning mode; the outlook's
|
||||
"installed base" framing inherits this caveat until a gate-plus-multiple-
|
||||
adapters (or interference-free training) configuration is shown. MBPP
|
||||
likely overlaps the base model's pretraining data; both arms share any
|
||||
contamination, and memorized items land in the easy bucket, so the hard
|
||||
bucket if anything over-represents genuinely novel problems — but bucket
|
||||
composition is contamination-sensitive. Sensitivity to α=0.3 and band width
|
||||
is unreported (the width-matched ablation arm partially addresses width).
|
||||
Adapter-only training may underestimate the ceiling (band-LoRA "rung 2"
|
||||
untested).
|
||||
One model family (gemma-4), two scales, three task families. Hard buckets
|
||||
are small (n=55/38/25); within-ladder orderings are not individually
|
||||
significant, and only the pooled hard effect and the HumanEval overall gain
|
||||
survive multiple-comparison scrutiny. Bucket membership derives from greedy
|
||||
labeling runs (consensus-k0 robustness check moves numbers <2 points, but
|
||||
both checks share the base model; an independent 12B-relabeling proxy is
|
||||
running). A third architecture family was not run; rung-2 was not run at
|
||||
12B. LiveCodeBench: single seed per arm, hard n=25, stdin-judged problems
|
||||
only, and its newest shard (Apr 2025) is *newer than MBPP by years* but
|
||||
not provably past the base model's undisclosed training cutoff — we claim
|
||||
recency, not proven non-contamination. The capability panel is
|
||||
MC-likelihood, not generation; its "no damage" answer does not extend to
|
||||
generation quality off-distribution (LCB shows trained arms *do* hurt
|
||||
there). The easy-item perturbation tax persists wherever the
|
||||
gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both
|
||||
arms share contamination, and memorized items land in the easy bucket, but
|
||||
bucket composition is contamination-sensitive. The capability panel
|
||||
(§3.8) is pending; until it lands, off-task effects of k>0 are unmeasured.
|
||||
The Blocksworld-12B and GSM8K-12B failures mean the adaptive-α fix is
|
||||
demonstrated on one task at one scale, not established as a general recipe.
|
||||
|
||||
## 6. Outlook
|
||||
## 6. Conclusion
|
||||
|
||||
The retrofit recipe — lens-locate, anchor-merge, verifier-filtered
|
||||
curriculum, gate — is scale-portable by construction: trainable mass is
|
||||
independent of base size, and prompt-side loops are prefill-shaped, so their
|
||||
economics *improve* with scale while serial CoT decode gets slower. The open
|
||||
question that decides whether this is a curiosity or a method is whether the
|
||||
effect survives scale (12B next; then a mid-size uptraining of the band
|
||||
itself). If it does, "loopification" becomes a cheap post-training phase any
|
||||
holder of a pretrained model can apply — a silent planning mode for the
|
||||
installed base, with its latent reasoning legible to the same lens that
|
||||
built it.
|
||||
The experiment this program set out to run — *can an interpretability lens
|
||||
tell you where to install recurrence in a frozen model, and does it work?* —
|
||||
has a clean answer: yes, and the placement is causally load-bearing. The
|
||||
more interesting answer is what the recurrence turned out to be: not a
|
||||
reasoning engine, but a remarkably cheap way to make a frozen model amortize
|
||||
its own planning into 0.03% of extra parameters, with a training-time loop
|
||||
as scaffold and an inference-time loop that is optional once the content
|
||||
exists. The practical recipe that survives all controls: lens-locate the
|
||||
band; anchor-merge with a state-dependent coefficient; label difficulty by
|
||||
STaR; distill plans if you have them, loop if you don't; gate by predicted
|
||||
difficulty; keep k=0 as the exact base model. What it buys: the
|
||||
plan-dependent slice at zero tokens and zero decode cost. What it does not
|
||||
buy: overall accuracy beyond what matched-compute sampling already delivers.
|
||||
Both halves of that sentence are the contribution.
|
||||
|
||||
## Appendix A: compute accounting (FLOPs / wall-clock / tokens, separated)
|
||||
|
||||
Let P = prompt tokens, G = generated tokens, c = FLOPs per token per full
|
||||
forward pass. The band is 17 of 35 decoder layers at E2B (fraction
|
||||
f≈0.486) and 10 of 48 at 12B (f≈0.208).
|
||||
|
||||
**Latent loop, k=4, prompt-only.** Prefill: 1 base pass + 4 band passes
|
||||
over prompt positions = (1+4f)·cP ≈ **2.94·cP** at E2B (1.83× at 12B —
|
||||
the overhead *shrinks* with scale because lens bands grow sublinearly).
|
||||
Decode: exactly cG (looped states written to the KV cache once;
|
||||
bit-exactness verified). Wall-clock: prefill is compute-bound and
|
||||
position-parallel, but the k iterations are serial — prefill latency
|
||||
≈2.9×, typically a small fraction of end-to-end latency for G≫0.
|
||||
Visible tokens: +0.
|
||||
|
||||
**Best-of-3 sampling.** FLOPs: with shared prompt prefill (favorable
|
||||
accounting), cP + 3·cG ≈ cP + 3cG; without sharing 3c(P+G). For MBPP
|
||||
(P≈150–300, G≈150–220), the *extra* FLOPs vs direct (≈2cG) are of the same
|
||||
order as the loop's extra (≈1.94cP) — hence "≈matched". Wall-clock: 3×G
|
||||
serial bandwidth-bound decode steps (or 3 parallel decode streams at 3×
|
||||
memory); strictly worse latency than the loop unless parallelized.
|
||||
Visible tokens: ≈3× (two discarded candidates). Requires a verifier or
|
||||
selector to pick among samples for the overall win we report (we use
|
||||
any-pass, an upper bound — see §3.4 caveat).
|
||||
|
||||
**Budget-CoT (50-token plan).** FLOPs: ≈c(P+G+50) plus the plan tokens'
|
||||
KV in context for the remainder — the *cheapest* arm in FLOPs. Wall-clock:
|
||||
+50 serial decode steps before answer tokens start (worst first-token
|
||||
latency). Visible tokens: +50.
|
||||
|
||||
Summary: no single scalar makes these three arms "equal"; the loop
|
||||
dominates on tokens and decode latency, budget-CoT on FLOPs, best-of-3 on
|
||||
overall accuracy. §3.4's "≈matched FLOPs" refers to the extra-FLOPs
|
||||
order-of-magnitude equivalence above, not exact equality; the honest
|
||||
statement is the three-way trade-off, and we report all three axes.
|
||||
|
||||
@@ -0,0 +1,131 @@
|
||||
# Prototype plan: the self-paced workspace (v2)
|
||||
|
||||
*Drafted 2026-07-16, pre-registration-style. Goal: test whether the model
|
||||
can learn to allocate workspace-loop compute ON ITS OWN — per prompt and
|
||||
per generation step — rather than at a swept hyperparameter k.*
|
||||
|
||||
## The concept
|
||||
|
||||
At every step the system chooses: emit, or spend a band iteration updating
|
||||
the workspace first. Make that choice a learned gate g(workspace state).
|
||||
Compute becomes a decision, not a constant. Gate-bought iterations emit no
|
||||
tokens, so they are exposure-safe by construction (deterministic given
|
||||
state — the pause-position property).
|
||||
|
||||
## What already exists (de-risked ingredients)
|
||||
|
||||
| ingredient | evidence | where |
|
||||
|---|---|---|
|
||||
| static per-prompt gate (E0) | gated 52.0 overall, easy 97.5 (vs 88.5 uniform), hard 28.6; bottleneck = probe recall (18/28 tp, 95 predicted hard) | `gate_probe.py`, `eval_gated.json` |
|
||||
| state-dependent control heads train | adaptive-α rescued 12B (3.8K params) | AdaptiveMergeAdapter |
|
||||
| evolving state during generation | carry beats registers on GSM (hard 0→9.4) | `carry_common.py`, `eval_carry.json` |
|
||||
| loop-capacity knob | rung-2 band-LoRA = best hard numbers (42.9/46.4) | `lora_band.py` |
|
||||
| gates over DEPTH (complementary axis) | learnable compute envelope g[t,l] | `path_gates.py` (Nils, in progress) |
|
||||
| free gate supervision | STaR difficulty labels; per-position labels derivable | prep_star/prep_mbpp |
|
||||
|
||||
## Experiments
|
||||
|
||||
### E1 — learned per-prompt halting (prompt side, MBPP)
|
||||
Replace fixed k with a trained soft halting gate. Architecture: after each
|
||||
iteration i, gate head h(e, ŝ_i) → p_halt,i (zero-init to fixed-k
|
||||
behavior); training uses the soft mixture of iteration outputs weighted by
|
||||
halting distribution (ACT-style), CE + λ·E[iterations] compute penalty;
|
||||
deploy = argmax halt. Trains end-to-end, NO RL. Arms: λ ∈ {1e-3, 1e-2},
|
||||
vs E0 probe-gate and uniform-k anchors.
|
||||
**Pre-registered predictions:** (a) accuracy ≥ uniform k=4 overall at ≤60%
|
||||
of its mean iterations; (b) easy ≥ 95% (gate protects the substrate);
|
||||
(c) allocation correlates with STaR label (point-biserial r > 0.3);
|
||||
(d) hard ≥ E0's 28.6% (learned gate beats frozen probe recall).
|
||||
**Failure mode to watch:** gate collapse (all-0/all-1) — mitigate with
|
||||
penalty warmup + entropy bonus; collapse at all λ falsifies E1.
|
||||
|
||||
### E2 — generation-side gating (GSM, gated carry)
|
||||
Substrate: design-C carry + short verified-CoT supervision (dense targets;
|
||||
harvest with "solve in ≤3 short steps", answer-verified). Gate per token
|
||||
step decides whether the carry state updates through the band or passes
|
||||
through: x_t = g·merge(e_t, s_{t−1}) + (1−g)·e_t, penalty λ·E[g].
|
||||
Anchors: carry-always, carry-never (same supervision).
|
||||
**Predictions:** (a) gate fires non-uniformly, concentrated near numeric/
|
||||
operator tokens (measurable); (b) accuracy ≥ carry-always (gating as
|
||||
protection); (c) easy-bucket damage < carry-always's (83→45 was the
|
||||
unprotected number). Hard-bucket *gain* over carry-always is hoped for,
|
||||
not predicted.
|
||||
|
||||
### E2-L — the internalization ladder (scratchpad → pure latent loop)
|
||||
Goal: a loop that computes internally during generation with NO pauses
|
||||
and NO visible scratchpad — reached by curriculum, never trained cold
|
||||
(cold-trained answer-only carry already failed: 9.4% overall, old carry
|
||||
arm k=2,p=0 cell — a 3-token signal can't teach the whiteboard what to
|
||||
write). Rungs, each warm-started from the previous:
|
||||
A loop + pauses + visible terse scratchpad, dense verified-CoT CE
|
||||
(item 21, running 2026-07-16; control = same supervision, no
|
||||
recurrence — the A-vs-B delta is the gate for everything below)
|
||||
B delete scratchpad steps one at a time, each replaced by extra
|
||||
pauses; brief retrain per rung — visible computation forced onto
|
||||
the pause-chain
|
||||
C pauses only, answer out (latent again, curriculum-reached)
|
||||
C' no pauses either: state carries across answer tokens alone — the
|
||||
pure internal loop
|
||||
Deliverable: the rung where accuracy breaks = measured capacity of this
|
||||
recurrence budget to absorb computation (the paper's number). Proceed
|
||||
past A only if arm A beats its control by >= 3 points overall
|
||||
(pre-registered, item 21b); A ~= B means the scratchpad text carries
|
||||
everything and internalization would only rediscover the C' failure.
|
||||
|
||||
### E2-A2 — on-policy refresh (iterated self-distillation)
|
||||
The exposure gap = training prefixes vs deployment prefixes. Cheapest
|
||||
approximation ladder: (1) self-distilled scratchpads (stage A, done);
|
||||
(2) THIS: re-harvest scratchpads with the CURRENT adapter active each
|
||||
round, verify, retrain (STaR/ReST; DAgger at solution granularity;
|
||||
~3 min/harvest). Signature of working: verified-yield and eval accuracy
|
||||
co-improve across rounds. Gated on the A-vs-B verdict.
|
||||
|
||||
### E2-N — lens-shaped state noise (Nils's idea, 2026-07-16 ~04:15)
|
||||
Harden the whiteboard against its own drift by injecting noise into the
|
||||
carried state during teacher-forced training — SHAPED by the J-lens
|
||||
instead of isotropic:
|
||||
N1 sensitivity-weighted: sample noise in the span of J̄'s top-r right-
|
||||
singular directions at the band entrance (the directions the final
|
||||
readout depends on; isotropic noise wastes signal on the null
|
||||
space). Cheap: jbar.pt exists; --lensnoise rank,scale flag.
|
||||
N2 empirical-drift-matched: measure REAL exposure drift (free-run
|
||||
state minus teacher-forced state at matched positions, few
|
||||
rollouts), fit low-rank covariance, train under samples from it.
|
||||
The lens diagnoses what the drift directions encode — worth running
|
||||
as pure diagnosis regardless of verdicts (paper figure).
|
||||
N3 concept-jitter: lens-read the carried concept (e.g. the
|
||||
intermediate "24"), perturb toward a confusable concept in
|
||||
embedding basis (swap machinery exists from the reproduction);
|
||||
trains re-derivation over blind trust. Most ambitious.
|
||||
Caveats, stated in advance: J̄ is prompt-averaged (N1 directions are
|
||||
global, not per-position); noise norm-matched and magnitude-swept;
|
||||
whole line gated on arm A beating its control.
|
||||
|
||||
### E3 — power knob (only if E1 or E2 shows clean gating)
|
||||
Warm-start rung-2 band-LoRA under the gate; joint fine-tune. Question: do
|
||||
gate-bought iterations do MORE per iteration with a trainable band?
|
||||
Metric: the internalization count (how many scratchpad steps can be
|
||||
removed post-hoc, E2 curriculum) as a function of LoRA rank.
|
||||
|
||||
### Lens verification (throughout — our home advantage)
|
||||
J-lens reads of gated vs ungated positions: do bought iterations sharpen
|
||||
task-relevant concepts at the positions where the gate fired? This is the
|
||||
mechanistic check that the gate allocates *meaningfully*, not just
|
||||
correlationally.
|
||||
|
||||
## Explicitly out of scope for the prototype
|
||||
Outcome-RL training of the gate (GRPO with compute price) — stage 2, only
|
||||
if E1–E3 show selective gating. 12B/scale transfer. Cross-task gates.
|
||||
|
||||
## Budget & order
|
||||
E1: 3 arms × ~75 min (Spark). E2: harvest ~30 min + 3 arms × ~90 min.
|
||||
E3: +2 arms. Total ≈ 1.5 Spark-days. Runs after the lens campaign; queue
|
||||
via gpuq as usual, every arm pre-registered in PROTOCOL_UNIFIED.md before
|
||||
launch (items 18+).
|
||||
|
||||
## Kill criteria (decided in advance)
|
||||
- E1 gate collapse at all λ AND E2 uniform firing → the state does not
|
||||
carry usable "needs compute" signal at this scale; program stops, E0's
|
||||
static-gate deployment note stands as the practical answer.
|
||||
- E1 works but hard < E0 → learned gate worse than probe; ship probe-gate,
|
||||
keep E2 only if its (a)/(b) hold.
|
||||
@@ -158,3 +158,11 @@ directly to Lys's distribution-shift account and predicts the queued
|
||||
auto-alignment (adaptive η, training-free) is the natural fallback if no
|
||||
fixed α transfers. Worth a paragraph in Paper A (design justification +
|
||||
12B analysis) and Paper B (stability mechanism).
|
||||
|
||||
**Outcome (2026-07-14, prediction confirmed):** the `12b_adaptive` MBPP
|
||||
arm (adaptive anchoring, 2×/8× H100 fleet) recovers the substrate —
|
||||
k=0/2/4 = 72.4/70.8/69.4 overall vs 72.6/54.6/43.0 under fixed α=0.3,
|
||||
with hard monotone 11.4→18.2→27.3. The tolerable-loop-share account
|
||||
called this shape in advance: stability restored by weakening the loop
|
||||
share, hard gains preserved and k-monotone, residual easy erosion
|
||||
(98.9→92.2) to be handled by the gate.
|
||||
|
||||
@@ -0,0 +1,78 @@
|
||||
# The full matrix: items 1–31 (as of 2026-07-17 late)
|
||||
|
||||
Scored source of truth: PROTOCOL_UNIFIED.md. All items pre-registered before running.
|
||||
|
||||
## Arc I — Band-loop retrofit (MBPP; frozen E2B + merge adapter)
|
||||
|
||||
| # | tried | key result | control/reference | verdict |
|
||||
|---|---|---|---|---|
|
||||
| 1–4 | unified merge adapter, k-loop over prompt | MBPP hard 3.6→28.6 (k=2); GSM hard 0→6.3 | k=0 same harness | loop works on code; GSM fails from day one |
|
||||
| 5 | same-size adapter, no recurrence | 17.9 hard | vs 28.6+ looped | loop > weights — recurrence load-bearing |
|
||||
| 6 | mixed-task training | both tasks regressed | single-task arms | interference, no synergy |
|
||||
| 7 | band-location ablation | L14–30: 43.6 vs early 23.6 / shifted 21.8 | width-matched | lens's "where" confirmed; some bands structurally null (KV-share) |
|
||||
| 9 | anchor sweep 13/12/11 | 34/31/25% overall, monotone collapse | L14 anchor | L14 boundary special |
|
||||
| 10 | L9 anchor (full-attention layer) | catastrophic (≤21.4 hard) | anchors 11–13 | lens boundary, not layer type |
|
||||
|
||||
## Arc II — Adapter-class factorial
|
||||
|
||||
| # | tried | key result | verdict |
|
||||
|---|---|---|---|
|
||||
| 11 | unconstrained RecurrentAdapter | ρ→4.5, easy 98→69, no depth gain | 4× params bought nothing |
|
||||
| 12 | Parcae (ρ<1 certified) | robust training, saturates, easy still 71% | stability ≠ fidelity — independent dials |
|
||||
| 13a | per-depth adapters (LTV) | hard content depth-stranded (17.9 @ k=2) | weight-sharing load-bearing |
|
||||
| 13b | free-ACT probe | no state fixed point; no difficulty gradient | output-stable orbit; no free halting |
|
||||
| 14 | tied-alpha (anchored B) | easy 93 preserved; α never moves | free B = fidelity culprit; 0.3 optimal |
|
||||
| 15 | randk / noise-s₀ / h2048 / seeds | 50–54 cells → seed means 37–46 | lucky seeds; only seed means are levels |
|
||||
| 16 | cross-task transfer | code adapter on GSM toxic (easy →28–45%) | content task-local, monotone ladder |
|
||||
|
||||
## Arc III — Gates & the GSM boundary
|
||||
|
||||
| # | tried | key result | verdict |
|
||||
|---|---|---|---|
|
||||
| 17 | GSM-only training, best recipe | hard ≤8.7 | structural → supervision-density diagnosis |
|
||||
| 18/19 | learned halting heads (E1a/b/c) | all lose to E0 frozen probe; E1c easy routing 95.9 @ 0.11 iters | classifier quality binds; hard recall regressed |
|
||||
| 20 | threshold curve + oracle | oracle 59.6 @ 0.24 iters; hards depth-diverse | gate worth ~9.6 pts, unclaimed |
|
||||
|
||||
## Arc IV — Hybrid & internalization ladder (GSM)
|
||||
|
||||
| # | tried | matched | ablated/control | verdict |
|
||||
|---|---|---|---|---|
|
||||
| 21 | carry + dense self-distilled scratchpads | **57.4** | FF control 54.7 | 5× prior best; recurrence edge = drop bucket, p=0.0094 |
|
||||
| 22 | delete steps d=1/2/3, +10 pauses each | 31.6 / 18.4 / 19.1 | base 10.9, cold 9.4 | breaks at d=1; plateau 2× cold (p=0.0025) |
|
||||
| 23a | 3× training steps | val ↑ (overfit) | — | time not binding |
|
||||
| 23b | 3× pauses | 29.3 | 31.6 | bandwidth not binding |
|
||||
| 24 | loop-only band-LoRA r16 | stopped (fit unchanged) | — | expressivity not binding; k=0 bit-exact validated |
|
||||
|
||||
## Arc V — Lens/state supervision of the latent chain
|
||||
|
||||
| # | tried | matched | ablated | verdict |
|
||||
|---|---|---|---|---|
|
||||
| 25 | lens-CE: pause j ↔ deleted token j | 31.2 | — | lce 10.3→1.9 yet flat: writing ≠ computing |
|
||||
| 26 | result-staging on pre-'=' spans | 24.2 / 15.6 combined | — | harmful — violates just-in-time schedule |
|
||||
| 27 | zero-pause 10-iter burst + lens | 34.0 | 32.4 (p=0.45) | best nominal, ns; pause tape dead weight |
|
||||
| 28 | teacher-state endpoint distillation | 30.1 | 29.7 | cos .113→.044, function absent; easy damaged |
|
||||
|
||||
## Arc VI — 2026-07-17 designs (Nils)
|
||||
|
||||
| # | tried | matched | ablated/ref | verdict |
|
||||
|---|---|---|---|---|
|
||||
| 29 | trajectory TF (10 waypoint transitions) | **39.1, p=0.045** | 39.1 burst-off (p=1.0) | first significant positive — a training signal, not an inference loop; +fr 30.5 (hurts); answer-only 15.6 < 19.1 |
|
||||
| 30 | metacog readiness head on carried state | AUC 0.798 | FF AUC 0.791 | signal real & cheap, NOT recurrence-specific; early-stop loses; oracle +3.9 |
|
||||
| 31 | synthetic KV memory (per-layer prefix) | 39.5 | 39.1 (p=1.0) | flat — gates frozen at −10 (init gradient-trap confound; −3 rerun open) |
|
||||
| 32 | discrete latent chain ("latent paper": lens-snapped symbols fed back) | TF 20.3 / ST 27.0 | 39.1 | net-harmful — exposure catastrophe (TF) and quantization noise (ST) both lose to the pure analog carry; architecture tree closed |
|
||||
|
||||
Instrument (unnumbered): whiteboard microscopy — three specimens + carry-vs-FF divergence (probe_discount*/probe_gsm*; board artifact).
|
||||
|
||||
## Standing positives
|
||||
MBPP loop-vs-weights gap · GSM drop-bucket reach (p=0.0094) · the 57.4 hybrid · trajectory-TF as a gradient (p=0.045) · 0.8-AUC readiness probe.
|
||||
|
||||
## Standing walls
|
||||
Consumption (8 write-side axes + 1 read-path attempt) · internalization (4 capacity axes + 4 supervision forms) · learned gates < frozen probe.
|
||||
|
||||
## Open threads
|
||||
Seeds for 39.1 · trajectory TF on rung A (move 57.4) · KV gate-init −3 · extension/deferral gating · g into the depth gate.
|
||||
|
||||
## Final architecture verdict (item 32 closes the tree)
|
||||
Pauses, bursts, KV memory, analog TF chains, and discrete chains all have controlled answers.
|
||||
The loop is a plan machine; tokens are the executor — they win by discreteness PLUS a verified
|
||||
commitment distribution (the LM head is trained to commit; the lens readout is not).
|
||||
@@ -0,0 +1,73 @@
|
||||
# Final statistics pass
|
||||
|
||||
## Headline numbers (Wilson 95% CIs)
|
||||
|
||||
- **MBPP untrained merge k=4 (n=500 rerun)**: overall 0.502 [0.458, 0.546] (n=500); hard 0.200 [0.116, 0.324] (n=55)
|
||||
- **MBPP trained FF k=1 (n=500 rerun)**: overall 0.536 [0.492, 0.579] (n=500); hard 0.273 [0.173, 0.402] (n=55)
|
||||
- **MBPP loop s0 k=0 (base)**: overall 0.518 [0.474, 0.561] (n=500); hard 0.055 [0.019, 0.149] (n=55)
|
||||
- **MBPP loop s0 k=2**: overall 0.530 [0.486, 0.573] (n=500); hard 0.309 [0.203, 0.440] (n=55)
|
||||
- **MBPP loop s0 k=4**: overall 0.536 [0.492, 0.579] (n=500); hard 0.436 [0.314, 0.567] (n=55)
|
||||
- **MBPP distill s1 k=1 (FF)**: overall 0.544 [0.500, 0.587] (n=500); hard 0.418 [0.297, 0.550] (n=55)
|
||||
- **MBPP pause16 k=1**: overall 0.552 [0.508, 0.595] (n=500); hard 0.364 [0.249, 0.496] (n=55)
|
||||
- **MBPP stack-train k=4**: overall 0.524 [0.480, 0.567] (n=500); hard 0.345 [0.234, 0.477] (n=55)
|
||||
- **MBPP distill-in-loopmode k=2**: overall 0.458 [0.415, 0.502] (n=500); hard 0.200 [0.116, 0.324] (n=55)
|
||||
- **Rust transfer k=0**: overall 0.591 [0.512, 0.665] (n=154); hard 0.080 [0.022, 0.250] (n=25)
|
||||
- **Rust transfer k=4**: overall 0.552 [0.473, 0.628] (n=154); hard 0.240 [0.115, 0.434] (n=25)
|
||||
- **HumanEval loop k=4**: overall 0.665 [0.589, 0.732] (n=164); hard 0.316 [0.191, 0.475] (n=38)
|
||||
- **HumanEval distill k=1**: overall 0.646 [0.571, 0.715] (n=164); hard 0.237 [0.130, 0.392] (n=38)
|
||||
- **MBPP best-of-3 (compute-matched)**: overall 0.572 [0.528, 0.615] (n=500); hard 0.327 [0.218, 0.459] (n=55)
|
||||
- **MBPP budget-CoT-50**: overall 0.538 [0.494, 0.581] (n=500); hard 0.400 [0.281, 0.532] (n=55)
|
||||
- **MBPP distill, 8 runs (hard)**: mean 0.457 ± 0.046 sd (range 0.400-0.545); overall mean 0.555
|
||||
- **MBPP loop seeds k=4 (hard)**: mean 0.375 ± 0.055 sd (n_seeds=5)
|
||||
|
||||
## McNemar exact tests (paired on items)
|
||||
|
||||
- loop k=4 vs k=0, overall: A-only 30, B-only 39, n=500, p=0.3356 (n.s.)
|
||||
- loop k=4 vs k=0, hard: A-only 1, B-only 22, n=55, p=5.722e-06 (**significant**)
|
||||
- loop k=4 vs UNTRAINED merge k=4, hard (net effect): A-only 4, B-only 17, n=55, p=0.007197 (**significant**)
|
||||
- loop k=4 vs UNTRAINED merge k=4, overall: A-only 29, B-only 46, n=500, p=0.06395 (n.s.)
|
||||
- trained FF vs UNTRAINED merge, hard: A-only 7, B-only 11, n=55, p=0.4807 (n.s.)
|
||||
- distill k=1 vs UNTRAINED merge k=4, hard: A-only 3, B-only 15, n=55, p=0.007538 (**significant**)
|
||||
- distill k=1 vs loop k=4, overall: A-only 33, B-only 37, n=500, p=0.7202 (n.s.)
|
||||
- distill k=1 vs loop k=4, hard: A-only 10, B-only 9, n=55, p=1 (n.s.)
|
||||
- stack-train k=4 vs distill k=1, hard: A-only 10, B-only 6, n=55, p=0.4545 (n.s.)
|
||||
- HumanEval loop k=4 vs k=0, overall: A-only 5, B-only 18, n=164, p=0.01062 (**significant**)
|
||||
- HumanEval distill k=1 vs k=0, overall: A-only 7, B-only 17, n=164, p=0.06391 (n.s.)
|
||||
- HumanEval trained vs UNTRAINED merge (k=2), overall: A-only 7, B-only 9, n=164, p=0.8036 (n.s.)
|
||||
- Rust loop k=4 vs k=0, overall: A-only 11, B-only 5, n=154, p=0.2101 (n.s.)
|
||||
- Rust loop k=4 vs k=0, hard: A-only 0, B-only 4, n=25, p=0.125 (n.s.)
|
||||
|
||||
## Token baselines (paired, per-item)
|
||||
|
||||
- **best-of-3 ORACLE (any-pass)**: overall 0.578 [0.534, 0.621] (n=500); hard 0.345 [0.234, 0.477] (n=55)
|
||||
- **best-of-3 oracle (selector run)**: overall 0.578 [0.534, 0.621] (n=500); hard 0.345 [0.234, 0.477] (n=55)
|
||||
- **best-of-3 DEPLOYABLE (logprob-selected)**: overall 0.550 [0.506, 0.593] (n=500); hard 0.273 [0.173, 0.402] (n=55)
|
||||
- bo3-oracle vs loop k=4, overall: arm-only 27, bo3-only 48, p=0.0203 (**significant**)
|
||||
- bo3-oracle vs loop k=4, hard: arm-only 14, bo3-only 9, p=0.4049 (n.s.)
|
||||
- bo3-oracle vs distill k=1, overall: arm-only 27, bo3-only 44, p=0.05681 (n.s.)
|
||||
- bo3-oracle vs distill k=1, hard: arm-only 14, bo3-only 10, p=0.5413 (n.s.)
|
||||
- bo3-deployable vs loop k=4, overall: arm-only 31, bo3-only 38, p=0.4704 (n.s.)
|
||||
- bo3-deployable vs loop k=4, hard: arm-only 16, bo3-only 7, p=0.09314 (n.s.)
|
||||
- bo3-deployable vs distill k=1, overall: arm-only 31, bo3-only 34, p=0.8043 (n.s.)
|
||||
- bo3-deployable vs distill k=1, hard: arm-only 15, bo3-only 7, p=0.1338 (n.s.)
|
||||
- **budget-CoT-50 (per-item rerun)**: overall 0.538 [0.494, 0.581] (n=500); hard 0.382 [0.265, 0.514] (n=55)
|
||||
- budget-CoT vs loop k=4, overall: arm-only 36, cot-only 37, p=1 (n.s.)
|
||||
- budget-CoT vs loop k=4, hard: arm-only 14, cot-only 11, p=0.69 (n.s.)
|
||||
- budget-CoT vs distill k=1, overall: arm-only 29, cot-only 26, p=0.7877 (n.s.)
|
||||
- budget-CoT vs distill k=1, hard: arm-only 8, cot-only 6, p=0.7905 (n.s.)
|
||||
|
||||
## Pooled hard bucket (MBPP + HumanEval + Rust)
|
||||
|
||||
Paired within-item k>0 vs k=0, counts pooled across benchmarks (loop arm; distill pooled where available).
|
||||
|
||||
- **loop**: base 5/118 -> loop 42/118 (0.042 -> 0.356, CI [0.275, 0.446]), McNemar p=1.46e-10
|
||||
- **distill**: base 3/93 -> distill 32/93 (0.032 -> 0.344, CI [0.255, 0.445]), McNemar p=2.98e-08
|
||||
|
||||
## Label robustness (consensus-k0 hard set)
|
||||
|
||||
Hard bucket redefined as: labeled hard AND k=0 fails in every seed's own eval run (removes single-greedy-run selection noise).
|
||||
|
||||
- consensus hard set: 52 of 55 labeled-hard items
|
||||
- MBPP loop s0 k=4: labeled-hard 0.436 -> consensus-hard 0.423 [0.299, 0.558] (n=52)
|
||||
- MBPP distill s1 k=1 (FF): labeled-hard 0.418 -> consensus-hard 0.404 [0.282, 0.539] (n=52)
|
||||
- MBPP stack-train k=4: labeled-hard 0.345 -> consensus-hard 0.346 [0.232, 0.482] (n=52)
|
||||
@@ -0,0 +1,47 @@
|
||||
{
|
||||
"tag": "rec16_e400_PARTIAL",
|
||||
"note": "k=16,32 cancelled by decision after k<=8 showed prediction (a); rho(A) from checkpoints",
|
||||
"rho_trajectory": {
|
||||
"e100": 3.357,
|
||||
"e200": 4.31,
|
||||
"e300": 4.456,
|
||||
"e400": 4.525,
|
||||
"e500": 4.492,
|
||||
"e600": 4.465
|
||||
},
|
||||
"ks": {
|
||||
"0": {
|
||||
"acc": 0.488,
|
||||
"by_label": {
|
||||
"easy": 0.984,
|
||||
"hard": 0.036,
|
||||
"drop": 0.01
|
||||
}
|
||||
},
|
||||
"2": {
|
||||
"acc": 0.408,
|
||||
"by_label": {
|
||||
"easy": 0.697,
|
||||
"hard": 0.357,
|
||||
"drop": 0.07
|
||||
}
|
||||
},
|
||||
"4": {
|
||||
"acc": 0.412,
|
||||
"by_label": {
|
||||
"easy": 0.697,
|
||||
"hard": 0.393,
|
||||
"drop": 0.07
|
||||
}
|
||||
},
|
||||
"8": {
|
||||
"acc": 0.408,
|
||||
"by_label": {
|
||||
"easy": 0.689,
|
||||
"hard": 0.429,
|
||||
"drop": 0.06
|
||||
}
|
||||
}
|
||||
},
|
||||
"n": 250
|
||||
}
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"tag": "base_k0",
|
||||
"k": 0,
|
||||
"arc": 0.36,
|
||||
"winogrande": 0.55875,
|
||||
"hellaswag": 0.5225,
|
||||
"mmlu": 0.30125
|
||||
}
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"tag": "distill_ff",
|
||||
"k": 1,
|
||||
"arc": 0.4175,
|
||||
"winogrande": 0.56625,
|
||||
"hellaswag": 0.57,
|
||||
"mmlu": 0.3175
|
||||
}
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"tag": "loop_k2",
|
||||
"k": 2,
|
||||
"arc": 0.36125,
|
||||
"winogrande": 0.5525,
|
||||
"hellaswag": 0.49625,
|
||||
"mmlu": 0.3125
|
||||
}
|
||||
@@ -0,0 +1,26 @@
|
||||
{
|
||||
"0": {
|
||||
"acc": 0.5,
|
||||
"by_label": {
|
||||
"easy": 0.9672131147540983,
|
||||
"hard": 0.17857142857142858,
|
||||
"drop": 0.02
|
||||
}
|
||||
},
|
||||
"2": {
|
||||
"acc": 0.504,
|
||||
"by_label": {
|
||||
"easy": 0.9098360655737705,
|
||||
"hard": 0.35714285714285715,
|
||||
"drop": 0.05
|
||||
}
|
||||
},
|
||||
"4": {
|
||||
"acc": 0.512,
|
||||
"by_label": {
|
||||
"easy": 0.9098360655737705,
|
||||
"hard": 0.42857142857142855,
|
||||
"drop": 0.05
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,26 @@
|
||||
{
|
||||
"0": {
|
||||
"acc": 0.496,
|
||||
"by_label": {
|
||||
"easy": 0.9672131147540983,
|
||||
"hard": 0.10714285714285714,
|
||||
"drop": 0.03
|
||||
}
|
||||
},
|
||||
"2": {
|
||||
"acc": 0.516,
|
||||
"by_label": {
|
||||
"easy": 0.9344262295081968,
|
||||
"hard": 0.39285714285714285,
|
||||
"drop": 0.04
|
||||
}
|
||||
},
|
||||
"4": {
|
||||
"acc": 0.524,
|
||||
"by_label": {
|
||||
"easy": 0.9098360655737705,
|
||||
"hard": 0.4642857142857143,
|
||||
"drop": 0.07
|
||||
}
|
||||
}
|
||||
}
|
||||
|
After Width: | Height: | Size: 202 KiB |
|
After Width: | Height: | Size: 84 KiB |
|
After Width: | Height: | Size: 104 KiB |
|
After Width: | Height: | Size: 171 KiB |
|
After Width: | Height: | Size: 66 KiB |
|
After Width: | Height: | Size: 80 KiB |
|
After Width: | Height: | Size: 43 KiB |
@@ -0,0 +1,93 @@
|
||||
{
|
||||
"curve": [
|
||||
{
|
||||
"theta": 0.3,
|
||||
"overall": 0.488,
|
||||
"by_label": {
|
||||
"easy": 0.9672131147540983,
|
||||
"hard": 0.07142857142857142,
|
||||
"drop": 0.02
|
||||
},
|
||||
"ek": 0.432
|
||||
},
|
||||
{
|
||||
"theta": 0.5,
|
||||
"overall": 0.5,
|
||||
"by_label": {
|
||||
"easy": 0.9590163934426229,
|
||||
"hard": 0.17857142857142858,
|
||||
"drop": 0.03
|
||||
},
|
||||
"ek": 0.74
|
||||
},
|
||||
{
|
||||
"theta": 0.7,
|
||||
"overall": 0.496,
|
||||
"by_label": {
|
||||
"easy": 0.9426229508196722,
|
||||
"hard": 0.17857142857142858,
|
||||
"drop": 0.04
|
||||
},
|
||||
"ek": 1.136
|
||||
},
|
||||
{
|
||||
"theta": 0.8,
|
||||
"overall": 0.508,
|
||||
"by_label": {
|
||||
"easy": 0.9344262295081968,
|
||||
"hard": 0.2857142857142857,
|
||||
"drop": 0.05
|
||||
},
|
||||
"ek": 1.388
|
||||
},
|
||||
{
|
||||
"theta": 0.9,
|
||||
"overall": 0.496,
|
||||
"by_label": {
|
||||
"easy": 0.8852459016393442,
|
||||
"hard": 0.39285714285714285,
|
||||
"drop": 0.05
|
||||
},
|
||||
"ek": 1.76
|
||||
},
|
||||
{
|
||||
"theta": 0.95,
|
||||
"overall": 0.492,
|
||||
"by_label": {
|
||||
"easy": 0.8770491803278688,
|
||||
"hard": 0.42857142857142855,
|
||||
"drop": 0.04
|
||||
},
|
||||
"ek": 2.084
|
||||
},
|
||||
{
|
||||
"theta": 0.98,
|
||||
"overall": 0.504,
|
||||
"by_label": {
|
||||
"easy": 0.8770491803278688,
|
||||
"hard": 0.4642857142857143,
|
||||
"drop": 0.06
|
||||
},
|
||||
"ek": 2.412
|
||||
},
|
||||
{
|
||||
"theta": 0.99,
|
||||
"overall": 0.508,
|
||||
"by_label": {
|
||||
"easy": 0.8770491803278688,
|
||||
"hard": 0.5,
|
||||
"drop": 0.06
|
||||
},
|
||||
"ek": 2.628
|
||||
}
|
||||
],
|
||||
"oracle": {
|
||||
"overall": 0.596,
|
||||
"by_label": {
|
||||
"easy": 1.0,
|
||||
"hard": 0.6428571428571429,
|
||||
"drop": 0.09
|
||||
},
|
||||
"ek": 0.236
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,152 @@
|
||||
{
|
||||
"abc382_a": true,
|
||||
"abc382_b": true,
|
||||
"abc382_c": false,
|
||||
"abc382_d": true,
|
||||
"abc382_f": false,
|
||||
"abc382_g": false,
|
||||
"abc383_a": true,
|
||||
"abc383_b": true,
|
||||
"abc383_c": false,
|
||||
"abc383_d": true,
|
||||
"abc383_e": false,
|
||||
"abc384_a": true,
|
||||
"abc384_b": true,
|
||||
"abc384_c": true,
|
||||
"abc384_d": false,
|
||||
"abc384_e": false,
|
||||
"abc384_f": false,
|
||||
"abc384_g": false,
|
||||
"abc385_a": true,
|
||||
"abc385_b": true,
|
||||
"abc385_c": false,
|
||||
"abc385_d": false,
|
||||
"abc385_e": false,
|
||||
"abc385_f": false,
|
||||
"abc386_a": false,
|
||||
"abc386_b": true,
|
||||
"abc386_c": false,
|
||||
"abc386_d": false,
|
||||
"abc386_e": false,
|
||||
"abc386_f": false,
|
||||
"abc387_a": true,
|
||||
"abc387_b": true,
|
||||
"abc387_c": false,
|
||||
"abc387_f": false,
|
||||
"abc388_a": true,
|
||||
"abc388_b": true,
|
||||
"abc388_c": false,
|
||||
"abc388_d": false,
|
||||
"abc388_e": false,
|
||||
"abc388_f": false,
|
||||
"abc388_g": false,
|
||||
"abc389_a": true,
|
||||
"abc389_b": true,
|
||||
"abc389_d": false,
|
||||
"abc389_e": false,
|
||||
"abc389_f": false,
|
||||
"abc389_g": false,
|
||||
"abc390_a": true,
|
||||
"abc390_b": true,
|
||||
"abc390_c": false,
|
||||
"abc390_d": false,
|
||||
"abc390_e": false,
|
||||
"abc390_f": false,
|
||||
"abc390_g": false,
|
||||
"abc391_a": true,
|
||||
"abc391_b": true,
|
||||
"abc391_d": false,
|
||||
"abc391_e": false,
|
||||
"abc391_f": false,
|
||||
"abc391_g": false,
|
||||
"abc392_a": true,
|
||||
"abc392_b": true,
|
||||
"abc392_c": false,
|
||||
"abc392_d": false,
|
||||
"abc392_f": true,
|
||||
"abc392_g": false,
|
||||
"abc393_a": true,
|
||||
"abc393_b": true,
|
||||
"abc393_d": false,
|
||||
"abc393_e": true,
|
||||
"abc393_f": false,
|
||||
"abc394_a": true,
|
||||
"abc394_b": true,
|
||||
"abc394_c": true,
|
||||
"abc394_d": true,
|
||||
"abc394_e": false,
|
||||
"abc394_f": false,
|
||||
"abc394_g": false,
|
||||
"abc395_a": true,
|
||||
"abc395_b": true,
|
||||
"abc395_c": true,
|
||||
"abc395_e": true,
|
||||
"abc395_f": false,
|
||||
"abc396_a": true,
|
||||
"abc396_b": true,
|
||||
"abc396_c": false,
|
||||
"abc396_d": true,
|
||||
"abc396_e": false,
|
||||
"abc396_f": false,
|
||||
"abc396_g": false,
|
||||
"abc397_a": true,
|
||||
"abc397_b": false,
|
||||
"abc397_c": true,
|
||||
"abc397_d": false,
|
||||
"abc397_e": false,
|
||||
"abc397_f": false,
|
||||
"abc397_g": false,
|
||||
"abc398_a": true,
|
||||
"abc398_b": false,
|
||||
"abc398_c": true,
|
||||
"abc398_d": false,
|
||||
"abc398_f": true,
|
||||
"abc398_g": false,
|
||||
"abc399_a": true,
|
||||
"abc399_b": true,
|
||||
"abc399_c": true,
|
||||
"abc399_d": false,
|
||||
"abc399_e": false,
|
||||
"abc399_f": false,
|
||||
"abc400_a": true,
|
||||
"abc400_b": true,
|
||||
"abc400_c": false,
|
||||
"abc400_d": false,
|
||||
"abc400_e": false,
|
||||
"abc400_g": false,
|
||||
"arc188_a": false,
|
||||
"arc188_b": false,
|
||||
"arc188_c": false,
|
||||
"arc188_d": false,
|
||||
"arc189_a": false,
|
||||
"arc189_b": false,
|
||||
"arc189_c": false,
|
||||
"arc189_d": false,
|
||||
"arc190_a": false,
|
||||
"arc190_c": false,
|
||||
"arc190_d": false,
|
||||
"arc191_a": false,
|
||||
"arc191_c": false,
|
||||
"arc191_d": true,
|
||||
"arc192_a": false,
|
||||
"arc192_b": false,
|
||||
"arc192_d": false,
|
||||
"arc192_e": false,
|
||||
"arc193_a": false,
|
||||
"arc193_b": false,
|
||||
"arc193_d": false,
|
||||
"arc194_a": false,
|
||||
"arc194_b": false,
|
||||
"arc194_c": false,
|
||||
"arc194_d": false,
|
||||
"arc194_e": false,
|
||||
"arc195_a": false,
|
||||
"arc195_b": false,
|
||||
"arc195_c": false,
|
||||
"arc195_d": false,
|
||||
"arc195_e": false,
|
||||
"arc196_a": false,
|
||||
"arc196_b": false,
|
||||
"arc196_c": false,
|
||||
"arc196_d": false
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
{"carry": 0.7981836131859141, "ff": 0.791359567919433}
|
||||
@@ -0,0 +1,561 @@
|
||||
{
|
||||
"MBPP untrained merge k=4 (n=500 rerun)": {
|
||||
"overall": {
|
||||
"acc": 0.502,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.458,
|
||||
0.546
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.2,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.116,
|
||||
0.324
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP trained FF k=1 (n=500 rerun)": {
|
||||
"overall": {
|
||||
"acc": 0.536,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.492,
|
||||
0.579
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.2727272727272727,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.173,
|
||||
0.402
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP loop s0 k=0 (base)": {
|
||||
"overall": {
|
||||
"acc": 0.518,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.474,
|
||||
0.561
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.05454545454545454,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.019,
|
||||
0.149
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP loop s0 k=2": {
|
||||
"overall": {
|
||||
"acc": 0.53,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.486,
|
||||
0.573
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.3090909090909091,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.203,
|
||||
0.44
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP loop s0 k=4": {
|
||||
"overall": {
|
||||
"acc": 0.536,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.492,
|
||||
0.579
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.43636363636363634,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.314,
|
||||
0.567
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP distill s1 k=1 (FF)": {
|
||||
"overall": {
|
||||
"acc": 0.544,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.5,
|
||||
0.587
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.41818181818181815,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.297,
|
||||
0.55
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP pause16 k=1": {
|
||||
"overall": {
|
||||
"acc": 0.552,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.508,
|
||||
0.595
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.36363636363636365,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.249,
|
||||
0.496
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP stack-train k=4": {
|
||||
"overall": {
|
||||
"acc": 0.524,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.48,
|
||||
0.567
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.34545454545454546,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.234,
|
||||
0.477
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP distill-in-loopmode k=2": {
|
||||
"overall": {
|
||||
"acc": 0.458,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.415,
|
||||
0.502
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.2,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.116,
|
||||
0.324
|
||||
]
|
||||
}
|
||||
},
|
||||
"Rust transfer k=0": {
|
||||
"overall": {
|
||||
"acc": 0.5909090909090909,
|
||||
"n": 154,
|
||||
"ci": [
|
||||
0.512,
|
||||
0.665
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.08,
|
||||
"n": 25,
|
||||
"ci": [
|
||||
0.022,
|
||||
0.25
|
||||
]
|
||||
}
|
||||
},
|
||||
"Rust transfer k=4": {
|
||||
"overall": {
|
||||
"acc": 0.551948051948052,
|
||||
"n": 154,
|
||||
"ci": [
|
||||
0.473,
|
||||
0.628
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.24,
|
||||
"n": 25,
|
||||
"ci": [
|
||||
0.115,
|
||||
0.434
|
||||
]
|
||||
}
|
||||
},
|
||||
"HumanEval loop k=4": {
|
||||
"overall": {
|
||||
"acc": 0.6646341463414634,
|
||||
"n": 164,
|
||||
"ci": [
|
||||
0.589,
|
||||
0.732
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.3157894736842105,
|
||||
"n": 38,
|
||||
"ci": [
|
||||
0.191,
|
||||
0.475
|
||||
]
|
||||
}
|
||||
},
|
||||
"HumanEval distill k=1": {
|
||||
"overall": {
|
||||
"acc": 0.6463414634146342,
|
||||
"n": 164,
|
||||
"ci": [
|
||||
0.571,
|
||||
0.715
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.23684210526315788,
|
||||
"n": 38,
|
||||
"ci": [
|
||||
0.13,
|
||||
0.392
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP best-of-3 (compute-matched)": {
|
||||
"overall": {
|
||||
"acc": 0.572,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.528,
|
||||
0.615
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.32727272727272727,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.218,
|
||||
0.459
|
||||
]
|
||||
}
|
||||
},
|
||||
"MBPP budget-CoT-50": {
|
||||
"overall": {
|
||||
"acc": 0.538,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.494,
|
||||
0.581
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.4,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.281,
|
||||
0.532
|
||||
]
|
||||
}
|
||||
},
|
||||
"distill_seed_spread": {
|
||||
"hard": [
|
||||
0.4909090909090909,
|
||||
0.41818181818181815,
|
||||
0.5454545454545454,
|
||||
0.43636363636363634,
|
||||
0.45454545454545453,
|
||||
0.4727272727272727,
|
||||
0.43636363636363634,
|
||||
0.4
|
||||
],
|
||||
"overall": [
|
||||
0.55,
|
||||
0.544,
|
||||
0.562,
|
||||
0.546,
|
||||
0.566,
|
||||
0.562,
|
||||
0.564,
|
||||
0.55
|
||||
]
|
||||
},
|
||||
"loop_seed_spread_hard": [
|
||||
0.43636363636363634,
|
||||
0.41818181818181815,
|
||||
0.3090909090909091,
|
||||
0.32727272727272727,
|
||||
0.38181818181818183
|
||||
],
|
||||
"mcnemar": {
|
||||
"loop k=4 vs k=0, overall": {
|
||||
"n": 500,
|
||||
"a_only": 30,
|
||||
"b_only": 39,
|
||||
"p": 0.33555761823401514
|
||||
},
|
||||
"loop k=4 vs k=0, hard": {
|
||||
"n": 55,
|
||||
"a_only": 1,
|
||||
"b_only": 22,
|
||||
"p": 5.7220458984375e-06
|
||||
},
|
||||
"loop k=4 vs UNTRAINED merge k=4, hard (net effect)": {
|
||||
"n": 55,
|
||||
"a_only": 4,
|
||||
"b_only": 17,
|
||||
"p": 0.007197380065917969
|
||||
},
|
||||
"loop k=4 vs UNTRAINED merge k=4, overall": {
|
||||
"n": 500,
|
||||
"a_only": 29,
|
||||
"b_only": 46,
|
||||
"p": 0.06394991646706696
|
||||
},
|
||||
"trained FF vs UNTRAINED merge, hard": {
|
||||
"n": 55,
|
||||
"a_only": 7,
|
||||
"b_only": 11,
|
||||
"p": 0.480682373046875
|
||||
},
|
||||
"distill k=1 vs UNTRAINED merge k=4, hard": {
|
||||
"n": 55,
|
||||
"a_only": 3,
|
||||
"b_only": 15,
|
||||
"p": 0.007537841796875
|
||||
},
|
||||
"distill k=1 vs loop k=4, overall": {
|
||||
"n": 500,
|
||||
"a_only": 33,
|
||||
"b_only": 37,
|
||||
"p": 0.7202027723528613
|
||||
},
|
||||
"distill k=1 vs loop k=4, hard": {
|
||||
"n": 55,
|
||||
"a_only": 10,
|
||||
"b_only": 9,
|
||||
"p": 1.0
|
||||
},
|
||||
"stack-train k=4 vs distill k=1, hard": {
|
||||
"n": 55,
|
||||
"a_only": 10,
|
||||
"b_only": 6,
|
||||
"p": 0.454498291015625
|
||||
},
|
||||
"HumanEval loop k=4 vs k=0, overall": {
|
||||
"n": 164,
|
||||
"a_only": 5,
|
||||
"b_only": 18,
|
||||
"p": 0.010622024536132812
|
||||
},
|
||||
"HumanEval distill k=1 vs k=0, overall": {
|
||||
"n": 164,
|
||||
"a_only": 7,
|
||||
"b_only": 17,
|
||||
"p": 0.06391465663909912
|
||||
},
|
||||
"HumanEval trained vs UNTRAINED merge (k=2), overall": {
|
||||
"n": 164,
|
||||
"a_only": 7,
|
||||
"b_only": 9,
|
||||
"p": 0.803619384765625
|
||||
},
|
||||
"Rust loop k=4 vs k=0, overall": {
|
||||
"n": 154,
|
||||
"a_only": 11,
|
||||
"b_only": 5,
|
||||
"p": 0.210113525390625
|
||||
},
|
||||
"Rust loop k=4 vs k=0, hard": {
|
||||
"n": 25,
|
||||
"a_only": 0,
|
||||
"b_only": 4,
|
||||
"p": 0.125
|
||||
},
|
||||
"bo3-oracle vs loop k=4, overall": {
|
||||
"n": 500,
|
||||
"a_only": 27,
|
||||
"b_only": 48,
|
||||
"p": 0.020298406990107445
|
||||
},
|
||||
"bo3-oracle vs loop k=4, hard": {
|
||||
"n": 55,
|
||||
"a_only": 14,
|
||||
"b_only": 9,
|
||||
"p": 0.4048728942871094
|
||||
},
|
||||
"bo3-oracle vs distill k=1, overall": {
|
||||
"n": 500,
|
||||
"a_only": 27,
|
||||
"b_only": 44,
|
||||
"p": 0.056814677932839015
|
||||
},
|
||||
"bo3-oracle vs distill k=1, hard": {
|
||||
"n": 55,
|
||||
"a_only": 14,
|
||||
"b_only": 10,
|
||||
"p": 0.5412561893463135
|
||||
},
|
||||
"bo3-deployable vs loop k=4, overall": {
|
||||
"n": 500,
|
||||
"a_only": 31,
|
||||
"b_only": 38,
|
||||
"p": 0.4703685318581444
|
||||
},
|
||||
"bo3-deployable vs loop k=4, hard": {
|
||||
"n": 55,
|
||||
"a_only": 16,
|
||||
"b_only": 7,
|
||||
"p": 0.0931396484375
|
||||
},
|
||||
"bo3-deployable vs distill k=1, overall": {
|
||||
"n": 500,
|
||||
"a_only": 31,
|
||||
"b_only": 34,
|
||||
"p": 0.8043170001933986
|
||||
},
|
||||
"bo3-deployable vs distill k=1, hard": {
|
||||
"n": 55,
|
||||
"a_only": 15,
|
||||
"b_only": 7,
|
||||
"p": 0.13380050659179688
|
||||
},
|
||||
"budget-cot vs loop k=4, overall": {
|
||||
"n": 500,
|
||||
"a_only": 36,
|
||||
"b_only": 37,
|
||||
"p": 1.0
|
||||
},
|
||||
"budget-cot vs loop k=4, hard": {
|
||||
"n": 55,
|
||||
"a_only": 14,
|
||||
"b_only": 11,
|
||||
"p": 0.6900379657745361
|
||||
},
|
||||
"budget-cot vs distill k=1, overall": {
|
||||
"n": 500,
|
||||
"a_only": 29,
|
||||
"b_only": 26,
|
||||
"p": 0.7877061896700435
|
||||
},
|
||||
"budget-cot vs distill k=1, hard": {
|
||||
"n": 55,
|
||||
"a_only": 8,
|
||||
"b_only": 6,
|
||||
"p": 0.79052734375
|
||||
}
|
||||
},
|
||||
"best-of-3 ORACLE (any-pass)": {
|
||||
"overall": {
|
||||
"acc": 0.578,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.534,
|
||||
0.621
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.34545454545454546,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.234,
|
||||
0.477
|
||||
]
|
||||
}
|
||||
},
|
||||
"best-of-3 oracle (selector run)": {
|
||||
"overall": {
|
||||
"acc": 0.578,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.534,
|
||||
0.621
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.34545454545454546,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.234,
|
||||
0.477
|
||||
]
|
||||
}
|
||||
},
|
||||
"best-of-3 DEPLOYABLE (logprob-selected)": {
|
||||
"overall": {
|
||||
"acc": 0.55,
|
||||
"n": 500,
|
||||
"ci": [
|
||||
0.506,
|
||||
0.593
|
||||
]
|
||||
},
|
||||
"hard": {
|
||||
"acc": 0.2727272727272727,
|
||||
"n": 55,
|
||||
"ci": [
|
||||
0.173,
|
||||
0.402
|
||||
]
|
||||
}
|
||||
},
|
||||
"pooled_hard_loop": {
|
||||
"base": 5,
|
||||
"arm": 42,
|
||||
"n": 118,
|
||||
"mcnemar": {
|
||||
"n": 118,
|
||||
"a_only": 1,
|
||||
"b_only": 38,
|
||||
"p": 1.4551915228366852e-10
|
||||
}
|
||||
},
|
||||
"pooled_hard_distill": {
|
||||
"base": 3,
|
||||
"arm": 32,
|
||||
"n": 93,
|
||||
"mcnemar": {
|
||||
"n": 93,
|
||||
"a_only": 1,
|
||||
"b_only": 30,
|
||||
"p": 2.9802322387695312e-08
|
||||
}
|
||||
},
|
||||
"consensus_hard": {
|
||||
"MBPP loop s0 k=4": {
|
||||
"acc": 0.4230769230769231,
|
||||
"n": 52,
|
||||
"ci": [
|
||||
0.299,
|
||||
0.558
|
||||
]
|
||||
},
|
||||
"MBPP distill s1 k=1 (FF)": {
|
||||
"acc": 0.40384615384615385,
|
||||
"n": 52,
|
||||
"ci": [
|
||||
0.282,
|
||||
0.539
|
||||
]
|
||||
},
|
||||
"MBPP stack-train k=4": {
|
||||
"acc": 0.34615384615384615,
|
||||
"n": 52,
|
||||
"ci": [
|
||||
0.232,
|
||||
0.482
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||