diff --git a/PAPER.md b/PAPER.md index d33bb14..dbd87f3 100644 --- a/PAPER.md +++ b/PAPER.md @@ -1,286 +1,353 @@ -# Retrofitting Latent Planning onto a Frozen Language Model via Workspace Recurrence +# Latent Planning by Workspace Recurrence: an Interpretability-Placed Implant, and What It Actually Buys -*Working draft, 2026-07-14. All experiments: google/gemma-4-E2B-it (frozen), single DGX Spark. Code and artifacts: `~/jspace`.* +*Final-data draft, 2026-07-14. Base models: google/gemma-4-E2B-it and +gemma-4-12B-it, both frozen. Hardware: DGX Spark + rented 2×/8×H100 nodes. +Code, per-item logs, and pre-registrations: `~/jspace` (git). Statistics: +`results-loop/STATS.md`.* ## Abstract -Interpretability work with an averaged-Jacobian lens ("J-lens") shows that -mid-depth layers of a pretrained language model form a *workspace*: a band of -layers that holds verbalizable, unspoken intermediate content. We ask whether -that band can be **iterated in place** — spending more serial compute per -input without emitting reasoning tokens — on a *frozen* model. A naive loop -diverges: the band is not a self-map. We show that a 1.6M-parameter -**anchor-dominant merge adapter** (0.03% of the model) at the band entrance -makes the recurrence a stable fixed-point iteration, and that training only -this adapter — with self-generated, verifier-filtered supervision and a -difficulty→depth curriculum — turns iteration into computation. On MBPP, -looping the workspace over the prompt ("latent planning") raises pass@1 on -plan-dependent problems from **5.5% to 30.9–43.6%** (three seeds, full test -set, execution-verified); overall accuracy is unchanged-to-slightly-improved -(51.8% → 51.8–53.8%, within noise at n=500) — the method's value is -cost-shaped (silent, prefill-parallel, no per-token overhead), not -accuracy-dominance. Controls attribute the hard-bucket gain to the -recurrence itself: a same-size adapter trained on identical data *without* -the loop reaches only 17.9%, exactly matching the untrained loop. On GSM8K the picture inverts — no recurrent variant beats the -weights-only control — and a four-arm decomposition localizes why: the loop -performs *plan refinement*, which code synthesis needs and answer-time -arithmetic does not. The J-lens provides both the intervention's design -(where to loop) and its verification (latent concepts sharpen ~8× per -converged iteration). Because the looped prompt states are constant during -generation, latent planning is prefill-shaped and adds no per-token cost. +Interpretability work with an averaged-Jacobian lens ("J-lens") partitions a +pretrained language model's depth into regimes, including a mid-depth +*workspace* band that holds verbalizable, unspoken intermediate content. We +retrofit recurrence onto this band in a **frozen** model: a 1.6M-parameter +anchor-dominant merge adapter (0.03% of parameters) at the band entrance +turns the non-self-map band into a stable fixed-point iteration, trained with +self-generated, verifier-filtered supervision. Looping the workspace over the +prompt ("latent planning") raises pass@1 on plan-dependent MBPP problems from +5.5% to 43.6% (seed mean 37.5±5.5), with zero visible tokens and zero +additional decode cost. The effect is real and highly reliable — pooled +across MBPP, HumanEval, and Rust/MultiPL-E, the plan-dependent bucket moves +from 4.2% to 35.6% (McNemar p≈1.5e-10) — and **placement is a law, not a +convenience**: the gain appears only when the loop enters at the +lens-identified boundary (L14), collapsing at L13 and below and structurally +nulling above. -## 1. Introduction +But a complete attribution program deflates the mechanism's mystique. +(i) **The loop's content is amortizable**: distilling the model's own +explicit plans into the same-size adapter — no recurrence at inference — +matches or exceeds the loop on the same bucket (mean over 8 runs 45.7±4.6 vs +37.5±5.5, paired difference n.s.), and the two do not stack; running the +loop on top of the distilled adapter *degrades* it. (ii) **Width rivals +depth**: 16 trained pause registers reach 36.4% on the same bucket. +(iii) **Compute-matched token baselines are uncomfortable**: best-of-3 +sampling beats every latent arm on overall accuracy (57.2% vs ≤55.2%), and a +50-token visible plan matches the loop on the hard bucket (40.0%). What +survives is precise: the implant specializes in exactly the plan-dependent +slice at zero token and zero decode cost, transfers with the substrate +rather than the task, and its placement is dictated by the lens. At 12B a +constant merge coefficient destroys the substrate; making the coefficient +state-dependent (a 3.8K-parameter gate) restores it on MBPP +(hard 11.4%→27.3% with overall preserved) but not on Blocksworld or GSM8K — +the stability dial that unifies this work with McLeish et al. (2511.07384) +and Lys et al. (2602.14759) is task- and scale-dependent. -Large language models buy reasoning accuracy with emitted tokens: chains of -thought give the network more serial passes, at the cost of latency, output -tokens, and bandwidth-bound decode. Recurrent-depth architectures (Universal -Transformers; DEQs; Huginn, arXiv:2502.05171; Mixture-of-Recursions, -arXiv:2507.10524) buy the same serial compute silently — but require -(pre)training the recurrence in at scale. +## 1. What this paper claims -We investigate a middle path: **retrofit** recurrence onto an off-the-shelf -frozen model, using an interpretability signal to decide *where*. The -J-lens (from the "verbalizable global workspace" line of work) partitions -depth into transduction, sensor, workspace, and motor regimes; the workspace -band (L14–30 of 35 in our subject model) holds slowly-varying, unspoken -intermediates — e.g. 'spider' before answering "8" to *"the animal that spins -webs has how many legs?"*. If the workspace approximates "iterate toward a -settled representation", looping it should deepen computation without -parameters. The contributions: +1. **A placement law.** The retrofit works if and only if the recurrence + enters at the lens boundary. Entrances at L9–L13 (same adapter, data, + curriculum) destroy overall accuracy (14–34% vs 52%) while recovering at + most half the hard-bucket gain; entrance at L14 preserves overall and + maximizes the gain (fig_placement). Entrances at L17/L24 are *structurally + null* in this architecture: KV-sharing makes layers ≥15 reuse keys/values + computed at ≤14, so k>0 is bit-identical to k=0 — a hazard for any + retrofit method that skips the mechanistic check. Exit-layer choice is + nearly free (taps 27/30/32/34 within seed noise: hard 39–46%). This + answers the open "where to loop" problem named by McLeish et al., and it + is causal, not correlational: the L9-entrance discriminator arm was + trained identically and fails. -1. **A minimal retrofit that works**: an anchor-dominant merge - (`(1−α)e + α·ŝ + MLP([e;ŝ])`, α=0.3, MLP zero-init, 1.6M params) makes - the frozen band a stable, answer-preserving recurrence; training only the - merge makes iterations *sharpen* rather than hold. -2. **A verified capability gain** on plan-dependent code synthesis, with the - full attribution grid (weights / untrained loop / trained loop / pause - tokens) showing the recurrence is the active ingredient. -3. **A mechanistic boundary**: math inverts the result, and the decomposition - (prompt-side vs generation-side × weights vs recurrence) identifies the - mechanism as plan refinement, not generic extra compute. -4. **Deployment properties**: bit-exact KV-cache-compatible inference (loop - once at prefill), a difficulty gate trained free from the labeling - pipeline, and economics that improve with model scale. +2. **A verified, statistically solid capability gain on a narrow slice.** + Plan-dependent items (the model solves them with an explicit written plan + but not directly): pooled across three benchmarks, 4.2%→35.6%, + p≈1.5e-10. Overall accuracy is statistically unchanged on MBPP + (p=0.34) and improved on HumanEval transfer (58.5%→66.5%, p=0.011). + +3. **A deflationary mechanism finding.** The trained loop converges to a + fixed point by k≈3–4 and behaves as *amortized plan content*, not + iterative computation: plan-distillation into the identical architecture + without recurrence matches it; stacking buys nothing (loop-training a + distill-warmed adapter: 34.5%, below distill alone; running the distilled + adapter in loop mode: drops to 20.0%); deeper k at inference is flat + (k=8: 40.0%). The recurrence is a *training-time scaffold* that lets the + adapter find plan-shaped content — content that can equally be put there + by distillation if plans are available. + +4. **A width-vs-depth law.** Trained pause registers (width) capture most of + the plan effect on code; recurrence (depth) is needed only where a state + must *evolve* — on GSM8K generation-side carry beats registers, and on + Blocksworld (pure planning, no world knowledge) the loop lifts hard-split + plans 0%→43% at 2B where everything else fails. Plans are wide; execution + is deep. + +5. **Honest economics.** The implant's costs: ≈2.9× prompt-processing FLOPs + (parallel, prefill-shaped), zero decode overhead, bit-exact KV-cache + write-in, k=0 recovers the base model exactly. Its competition at matched + FLOPs: best-of-3 sampling wins overall accuracy outright (57.2%); a + 50-token visible plan ties the hard bucket. The value proposition is + *only*: no visible tokens, no decode latency, and the hard-slice + specialization (distill's 46% > budget-CoT's 40% > best-of-3's 33%). + +6. **Scale transfers only with a state-dependent stability dial.** At 12B the + 2B-tuned constant α=0.3 collapses overall accuracy (72.6%→43.0%); the + damage is present *before* adapter training (untrained-loop arm) and is + not fixed by retuning α or LR. A per-position learned coefficient + α=σ(w·[e;ŝ]+b) restores MBPP (overall 69.4%, hard 11.4%→27.3%) — but + fails to rescue Blocksworld-12B and yields only a marginal GSM8K-12B + overall gain (35.9%→36.7% at k=1), the project's only overall 12B win. ## 2. Method **Locating the band.** The lens reads residual state h at layer ℓ through the -averaged Jacobian J̄_ℓ = E[∂h_final/∂h_ℓ] and the unembedding. Depth regimes -follow from what the readout tracks (input echo / abstract content / output -token). On gemma-4-E2B: workspace ≈ L14–30 (1.08B params, 58% of decoder). +averaged Jacobian J̄_ℓ = E[∂h_final/∂h_ℓ] and the unembedding; depth regimes +follow from what the readout tracks. On gemma-4-E2B: workspace ≈ L14–30 of +35; on 12B: L36–45 of 48. -**Making the band a self-map.** Feeding L30's output to L14 collapses in one -step (out-space ≠ in-space; norms and content 17 layers "downstream"). -Additive anchoring diverges. The fix is DEQ-style input injection done by -hand: with e = L13's output (fixed anchor) and s the fed-back band output, +**Making the band a self-map.** Feeding L30's output back to L14 collapses +(out-space ≠ in-space). With e = L13's output (fixed anchor) and s the +fed-back, norm-matched band output: - L14-in = (1−α)·e + α·(s · |e|/|s|) + MLP([e ; s·|e|/|s|]), α = 0.3. + L14-in = (1−α)·e + α·ŝ + MLP([e ; ŝ]), ŝ = s·|e|/|s| -Zero-initializing the MLP's output layer makes the untrained adapter exactly -the hand merge, which is stable and answer-preserving for ≥11 iterations but -only *holds* content (lens concept flat). +α=0.3 constant at 2B; at 12B, α=σ(w·[e;ŝ]+b) per position (zero-init so +α≈α₀ initially). MLP output zero-init: the untrained adapter is exactly the +hand merge — stable, answer-preserving, content-holding. -**Training only the merge.** Supervision is self-generated and -verifier-filtered (STaR-style): the frozen model attempts each task directly -and with explicit planning/CoT; items it solves only with planning are -"hard", direct solves "easy", neither "drop". Cross-entropy on answer/code -tokens of the *direct* prompt; the model's own verified outputs are the -targets (in-distribution). A **difficulty→depth curriculum** trains easy -items at loop depth k=1, mixed at k=2, hard only at k=2–4, so loss on hard -items is reducible only through the recurrence. For generation tasks the -loop applies to the **prompt span only** ("latent planning"): generated -tokens run the plain path but attend to the looped prompt states; this -removes exposure bias structurally. +**Training.** STaR-style self-labeling: items the frozen model solves only +with an explicit plan/CoT are "hard", direct solves "easy", neither "drop". +Cross-entropy on answer/code tokens of the direct prompt, model's own +verified outputs as targets. Difficulty→depth curriculum (easy k=1, mixed +k=2, hard k=2–4). The loop applies to the **prompt span only**; generated +tokens run the plain path but attend to looped prompt states. Variants +trained the same way: **pause-N** (N trained register tokens appended to the +prompt, no recurrence), **plan-distill** (KL from the model's own +plan-in-context distribution into the FF adapter), **rung-2** (warm-started +adapter + entrance-faded LoRA rank 8 on the band's first layers, loop-only +via a global toggle), and **stack** arms (distill-warm + loop training; +distilled adapter evaluated in loop mode). -**Inference cost.** Causality makes the looped prompt states independent of -generated tokens, so they are computed once; a hooked prefill writes them -into the KV cache and generation proceeds natively (verified bit-identical; -≥3.5× faster than recomputation). The concrete overhead at k=4 is 5 passes -over the band's 17/35 layers at prefill — ≈2.9× prompt-processing FLOPs, -parallel across positions — and **zero** additional decode cost. Explicit -planning with ~200 emitted tokens costs more total FLOPs and pays them -serially at bandwidth-bound decode; this asymmetry grows with model size. +**Inference.** Looped prompt states are causally independent of generated +tokens: computed once at prefill, written into the KV cache by a hooked +forward pass, generation native. Verified bit-identical to the slow path. +Cost at k=4: ≈2.9× prefill FLOPs, **zero** decode overhead. ## 3. Results -### 3.1 Latent planning on code (MBPP) +Statistics throughout: Wilson 95% CIs; paired comparisons by exact McNemar; +all headline arms evaluated on the full 500-item MBPP test split (hard +bucket n=55), HumanEval n=164 (hard n=38), Rust/MultiPL-E n=154 (hard n=25), +execution-verified. Label robustness: redefining "hard" as +labeled-hard ∧ k=0-fails-in-all-five-seeds (52/55 items) moves headline +numbers <2 points. -Full 500-item test split, greedy decode, unit-test-verified. Hard bucket = -items the frozen model solves only with an explicit written plan (n=55). +### 3.1 The placement law -| k=4 (prompt-only loops) | hard pass@1 | overall | +![Placement cliff](results-loop/fig_placement.png) + +Entrance-layer sweep with everything else fixed. L14 (lens boundary): +hard 43.6%, overall 53.6%. L13: hard 17.9%, overall 34.4%. L9–L12: overall +14.0–30.8% (substrate destroyed). L17/L24 entrances: k>0 ≡ k=0 (KV sharing; +verified bit-identical) — the 12B model has no shared-KV layers, making it +the unconfounded replication. Exit sweep at fixed entrance (L27/30/32/34): +hard 39.3–46.4%, within seed spread. The lens boundary is necessary; the +exit is a free parameter. (The L23-exit arm died in training and was not +rerun; the four completed exits bracket it.) + +### 3.2 The attribution ladder + +![Attribution ladder](results-loop/fig_ladder.png) + +MBPP hard bucket (plan-dependent, n=55 unless noted): + +| arm | hard pass@1 | overall | |---|---|---| -| baseline (k=0) | 5.5% | 51.8% | -| trained loop, seed 0 | **43.6%** | 53.6% | -| trained loop, seed 1 | **41.8%** | 53.8% | -| trained loop, seed 2 | **30.9%** | 51.8% | +| base (k=0, bit-exact) | 5.5% | 51.8% | +| untrained loop (α-merge only, n=28) | 17.9% | ~52% | +| trained FF, no recurrence (n=28) | 17.9% | ~52% | +| pause-16 registers (width) | 36.4% | 55.2% | +| **trained loop k=4** (seed mean, 5 seeds) | **37.5±5.5** (best 43.6) | 53.6% | +| rung-2: + entrance-faded band LoRA (n=28) | 42.9/46.4 (2 seeds) | 51.2/52.4 | +| **plan-distilled FF** (mean, 8 runs) | **45.7±4.6** (best 49.1) | 55.5% | +| budget-CoT (50 visible tokens) | 40.0% | 53.8% | +| best-of-3 sampling (≈matched FLOPs) | 32.7% | **57.2%** | +| explicit plan in context (ceiling) | 94.5% | 59.0% | -Silent loops recover roughly 40% of what explicit planning achieves, at zero -visible-token cost, with no overall regression (the easy-item perturbation -tax, ~9 points, is offset by hard/drop gains; a gate removes most of it, -§3.4). +Significance structure (McNemar, `STATS.md`): loop vs base on hard, +p=5.7e-6; every latent-arm-vs-latent-arm difference (loop vs distill, distill +vs stack) is **not significant** at n=55; loop vs base *overall* is not +significant on MBPP (p=0.34). The ladder's shape is reliable; its fine +ordering is not. -![MBPP pass@1 vs loop depth](results-loop/loop_eval_code.png) +### 3.3 The decisive tests: nothing stacks -### 3.2 Attribution: the recurrence is the ingredient +If the loop performed genuine iterative computation, plan-distilled content +plus recurrence should compound. It does not: -250-item subset; same data, same 1.6M parameters, same insertion point: +- **Distill-warm + loop training**: hard 34.5% — below distill alone. +- **Distilled adapter run in loop mode**: hard 20.0%, overall 45.8% — + looping *degrades* the distilled weights. +- **Pause-16 + distill**: hard 30.9% — no width stacking either. +- **Inference depth beyond convergence**: k=8 hard 40.0% ≈ k=4 (fixed point, + cos(sₖ,sₖ₋₁)=1.000 by k≈3–4). -| arm | hard pass@1 | -|---|---| -| baseline | 3.6% | -| untrained loop (α-merge only) | 17.9% | -| trained adapter, **no loop** (weights control) | 17.9% | -| trained **loop** | **42.9–46.4%** | +Reading: the recurrence is a **training-time scaffold**. The curriculum +forces hard-item loss to be reducible only through the loop, and what the +adapter learns to inject is plan-shaped content — the same content +distillation installs directly when explicit plans are available. The loop's +distinctive value is that it finds this content *without* plan supervision +(STaR labels only say which items needed plans, not what the plans were). -The weights control lands exactly on the untrained-loop value: ~18 points is -what perturbation-plus-format-alignment buys. The remaining ~28 points -require iterating the band. Post-hoc depth selection is excluded by -pre-registration (k=2 fixed on validation before test numbers existed; -k-curves reported descriptively). +### 3.4 Compute-matched honesty -**Checkpoint selection.** No checkpoint was chosen using test or generation -results. Seed 0's checkpoint (step 399) was fixed at training time from the -validation-CE overfitting inflection, before any generation eval of that -adapter; seeds 1–5 use step 400 by pre-commitment made before those seeds -were trained. We separately report that validation CE is a poor proxy for -generation accuracy (a checkpoint selected by val-CE on a sibling arm -underperformed a later one), which is why the fixed-step rule is used -rather than per-seed val selection. +At approximately matched FLOPs, token-space baselines are strong: best-of-3 +sampling wins overall accuracy against every latent arm (57.2%, +CI [52.8, 61.5], vs loop 53.6 [49.2, 57.9] — point estimate higher, CIs +overlap) by preserving easy items perfectly while sampling rescues some hard +ones. A 50-token visible plan ties the loop's hard bucket. The latent +implant's surviving advantages are qualitative: zero visible tokens (silent), +zero decode overhead (prefill-parallel; sampling and CoT pay serially at +bandwidth-bound decode), and the hard-slice crown under distillation (46% vs +40% budget-CoT vs 33% best-of-3). For deployment this means: the implant is +a *latency/token-budget* technology with a side specialization in +plan-dependent items — not an accuracy technology. -### 3.3 The boundary: math +### 3.5 Width vs depth, and the task boundary -On GSM8K, *no* recurrent variant beats the weights-only control. The four-arm -grid (hard bucket) decomposes the failure: +Pause registers (width) reach 36.4% (16 registers; 8: 30.9%, 32: 34.5% — flat +in N) on MBPP hard: static plan content fits in registers. GSM8K inverts the +prompt-side result entirely (no variant beats the weights control +prompt-side), but generation-side *carry* — recurrence across token steps — +doubles the pause control on hard items: arithmetic's serial state evolves +during the answer. Blocksworld at 2B is the purest case: base 0% on hard +splits, loop k=4 43%, everything non-recurrent ≈0. The law: **plans are +wide; execution is deep.** Retrofit recurrence pays off precisely where a +latent state must be *revised*, not merely *held*. -| GSM8K hard | prompt-side only | touches generation | -|---|---|---| -| feedforward weights | **11.8%** | 4.7% (pause-token control) | -| recurrence | 6.3–8.7% (prompt loop) | 9.4% (cross-token carry) | +### 3.6 Scale: the stability dial -Orthogonal effects: perturbing free-running generation positions is costly -for either mechanism; recurrence beats weights only where a state must -evolve (the generation side — carry doubles the pause control in-harness), -and loses on the static prompt side. No variant beats the 10.5% overall -baseline. Reading: the trained loop performs **plan refinement**; code -synthesis is plan-shaped, multi-step arithmetic is not — its serial -computation happens during the answer, and one frozen band pass per token -cannot perform it silently at 2B. CoT tokens remain load-bearing for math. -(Hard-bucket cells carry an outcome-selection caveat — buckets were defined -by greedy baseline outcomes; sampled relabeling is in progress — so the math -conclusion is stated on overall numbers.) +![Cross-scale grid](results-loop/fig_scale.png) -### 3.4 Mechanism and deployment +At 12B (no shared KV — unconfounded), constant α=0.3: overall collapses +72.6%→43.0% at k=4 while hard limps to 11.4%. The untrained-loop arm shows +the damage precedes adapter training; α=0.15 and LR retuning do not fix it +(47.6/52.6% overall). The state-dependent coefficient does, on MBPP: +overall 69.4% (base 72.4%), hard 11.4%→27.3%. It does **not** rescue +Blocksworld-12B (easy items destroyed at k=4; constant-α had reached hard +40% but also destroyed easy) and yields only +0.8 points overall on +GSM8K-12B (35.9→36.7 at k=1, hard 1.6→10.6) — the sole overall-accuracy win +of the program, and a marginal one. Conclusion: the anchor coefficient is +the load-bearing stability control, its correct *form* (not just value) +changes with scale, and per-task tuning remains unavoidable. -**Fixed point.** The trained loop takes a large first step -(cos(s₁,s₀)=0.926 vs 0.977 untrained) and converges bit-exactly by k≈3–4 -(cos=1.000), where accuracy and lens-sharpening plateau — extra iterations -are no-ops, explaining the k-curve shape. +### 3.7 Transfer: substrate, not task -![Loop convergence dynamics](results-loop/loop_dynamics.png) +![Transfer panel](results-loop/fig_transfer.png) -**Lens verification.** P(latent concept) under the J-lens at the band exit -rises 0.015→0.13 across iterations after training (~8× the untrained -control, which only holds). The same lens that located the band verifies -that looping deepens its computation — and makes the silent reasoning -inspectable. +MBPP-trained implants applied unchanged: **HumanEval** overall 58.5%→66.5% +(loop k=4, p=0.011; hard 0→31.6%); notably the *untrained* merge already +reaches 64.6% and the transferred pause adapter 66.5% (hard 38.9%) — the +transfer is substrate-shaped (a generically useful perturbation+content +mode), not task-memorized. **Rust/MultiPL-E** (Python-trained, different +language, compile-run-verified): hard 8.0%→24.0% (p=0.125 at n=25 — +directionally consistent, underpowered). **Blocksworld** MBPP-transfer: +hard 0→14.3% (task-trained: 43%). Content transfers where the substrate's +plan-representation overlaps; task-specific training still dominates. -**Gate.** A logistic probe on the k=0 workspace state (supervised for free -by the STaR labels) routes prompts: predicted-easy at k=0, predicted-hard at -k=4. Result: overall equal to the best uniform depth with easy items fully -preserved (97.5% vs 98.4% baseline); probe precision (19% at 64% recall) is -the current ceiling. +### 3.8 Mechanism, verification, deployment -**Negative results with content.** Mixed-task (code+math) training regressed -both tasks versus dedicated adapters, despite indistinguishable validation -CE — cross-entropy parity does not predict generation parity. Validation-CE -checkpoint selection likewise failed to track generation accuracy. +The trained loop takes a large first step (cos(s₁,s₀)=0.926 vs 0.977 +untrained) and converges bit-exactly by k≈3–4; accuracy and lens-sharpening +plateau there. P(latent concept) under the J-lens at the band exit rises +0.015→0.13 across iterations (~8× the untrained hold) — the lens that placed +the implant also renders its silent content inspectable. The STaR labels +train a free difficulty gate (route predicted-hard to k=4, else k=0); +gate quality (19% precision at 64% recall) is the current ceiling on +removing the easy-item perturbation tax. k=0 is the exact base model by +construction — the implant is removable at token granularity. + +**General-capability panel** (ARC-Challenge, WinoGrande, HellaSwag, MMLU; +length-normalized MC scoring with the loop applied to the context span) is +running on the Spark; results will quantify what k>0 does to off-task +abilities. [PENDING — fill on completion.] + +### 3.9 Negative results with content + +Mixed-task (code+math) training regressed both tasks at equal validation CE +— CE parity does not predict generation parity, and validation-CE checkpoint +selection fails likewise (fixed-step pre-commitment used instead; no +checkpoint was selected on test or generation results). GSM8K distillation +collapsed to empty outputs twice (E2B first attempt, 12B) on 3-token targets +under KL-dominant loss; a CE-dominant retry at E2B trained but reached only +hard 4.7%. Plan-distillation on GSM8K underperforms its MBPP twin even when +training succeeds: consistent with §3.5, there is little static plan content +for math to amortize. ## 4. Related work -Two recent papers bracket this work. **McLeish et al. (arXiv:2511.07384)** -retrofit depth-recurrence into pretrained 1B models via layer surgery + -continued pretraining (~50B tokens, all parameters, Muon, recurrence -curriculum to r=32): the generic claims "retrofitted recurrence works and -beats the non-recurrent parent" and "pretrain-then-convert" are theirs, at -~5 orders of magnitude more training cost than ours. They name layer choice -as an open problem; our lens-derived band with its causal backing (anchor -cliff at L14, tap invariance, wrong-band ≈ 0, KV-sharing hazard) is a direct -answer to it. Unlike their surgery (which needs a healing phase), our k=0 -exactly recovers the base model. **Lys et al. (arXiv:2602.14759)** loop -frozen models training-free and show naive looping degrades (distribution -shift) while interpolating with the un-looped state rescues it — independent -convergent evidence for our anchor-dominant merge; their whole setting -corresponds to the untrained cell of our attribution table (17.9% hard = -our FF/untrained level), evaluated by likelihood rather than execution. +**McLeish et al. (arXiv:2511.07384)** retrofit depth-recurrence via layer +surgery + ~50B-token continued pretraining of all parameters; they name +layer choice as an open problem — §3.1 is a causal answer. Their surgery +needs a healing phase; our k=0 is exactly the base model. **Lys et al. +(arXiv:2602.14759)** loop frozen models training-free; their finding that +naive looping degrades while interpolation with the un-looped state rescues +it is independent convergent evidence for anchor-dominance, and their +setting is the untrained cell of our ladder (17.9%). -**One mechanism, three regimes.** All three works are variants of a single -design: mix the fed-back state with an anchor derived from the un-looped -computation. Lys et al.'s inference-time moving average η·h⁽⁰⁾+(1−η)·h⁽ᵗ⁾ is -an untrained anchor-mixing coefficient; our (1−α)e + α·ŝ + MLP([e;ŝ]) is its -trained analogue (fixed mix + learned correction); McLeish et al.'s -concatenated input injection is the fully learned limit, trained end-to-end. -The anchor coefficient is the stability dial of frozen-band looping: Lys -et al.'s naive-looping collapse is the zero-anchor (α→1) limit, their -regularization gains are the untrained anchored regime, and our 12B failure -at the 2B-tuned α=0.3 — with the untrained-substrate arm showing the damage -is pre-training-of-the-adapter — is the same dial mis-set at a new scale. -Stability of retrofitted recurrence appears to be governed by how strongly -the loop is anchored, across all three training budgets. +**One mechanism, three regimes.** All three works mix the fed-back state +with an anchor from the un-looped computation. Lys et al.'s moving average +η·h⁽⁰⁾+(1−η)·h⁽ᵗ⁾ is an untrained anchor coefficient; our +(1−α)e + α·ŝ + MLP is its trained analogue; McLeish et al.'s input injection +is the fully-learned limit. The 12B episode closes the loop on this +unification: the coefficient is the stability dial, naive looping is its +α→1 collapse limit, and our scale failure + state-dependent fix show the +dial must itself become a function of the state as models grow. Our stacking +results add a caution for the whole family: if retrofitted recurrence +content is amortizable (§3.3), some of the family's gains may be +reproducible by distillation without inference-time recurrence — a control +neither bracket paper runs. -Earlier lineage: Universal Transformers (adaptive depth); DEQ (fixed-point -inference); Huginn (arXiv:2502.05171) — prelude/core/coda from scratch; -Mixture-of-Recursions (arXiv:2507.10524) — learned per-token depth; Relaxed -Recursive Transformers (arXiv:2410.20672) — uptrained tied layers; Coconut — -latent CoT; pause tokens (Goyal et al.) — token-space silent compute, whose -trained-adapter variant proved a near-match for our loop on MBPP (§3.2). +Earlier lineage: Universal Transformers; DEQ; Huginn (2502.05171); +Mixture-of-Recursions (2507.10524); Relaxed Recursive Transformers +(2410.20672); Coconut; pause tokens (Goyal et al.) — whose trained variant +proved a genuine rival, not a strawman (§3.2, §3.5). -What remains distinct here: **interpretability-derived loop placement with -causal validation** (answering McLeish et al.'s open problem); **a 1.6M-param -trained merge on a fully frozen base** (between Lys et al.'s free end and -McLeish et al.'s full-retraining end, and the only one of the three where -the base model is provably untouched); **prompt-only latent planning with -bit-exact KV-cache write-in and zero decode cost**; **the attribution -ladder** (untrained / weights / pause / loop / explicit plan) — neither -bracket paper runs compute-matched token-space controls; and **difficulty- -adaptive depth via the STaR-label gate**, named as future work in both. +What remains distinct here: interpretability-derived placement with causal +validation; a fully frozen base with bit-exact k=0 and zero-decode-cost KV +write-in; the complete attribution ladder including compute-matched +token-space baselines and stacking tests; the width/depth task law; and the +amortizability finding itself. ## 5. Limitations -One base model family at 2B-effective scale (12B replication in progress); -two task families. **Location specificity is not yet ablated**: a -pre-registered control looping shifted/early/late/width-matched bands with -identical adapter and curriculum is queued; until it lands, the results are -formally consistent with "any wide mid-depth band works", and the lens claim -rests on discovery convenience plus mechanism verification. Hard buckets are -small (n=55 greedy / n=33 sampled) with seed spread of ±6 items; sampled -relabeling shows 97% agreement with greedy labels, and intervals accompany -all bucket cells in the final tables. The MBPP attribution grid lacks a -pause-token arm and a plan-distillation baseline (both queued) — the GSM8K -grid has the former. Easy-item perturbation tax is not eliminated (gate -preserves easy items but probe precision is 19%). Visible planning remains -stronger on absolute accuracy — the claim is cost-and-latency-shaped. -**Mixed-task training regressed both tasks**, so the current recipe yields -per-task adapters, not one general silent-planning mode; the outlook's -"installed base" framing inherits this caveat until a gate-plus-multiple- -adapters (or interference-free training) configuration is shown. MBPP -likely overlaps the base model's pretraining data; both arms share any -contamination, and memorized items land in the easy bucket, so the hard -bucket if anything over-represents genuinely novel problems — but bucket -composition is contamination-sensitive. Sensitivity to α=0.3 and band width -is unreported (the width-matched ablation arm partially addresses width). -Adapter-only training may underestimate the ceiling (band-LoRA "rung 2" -untested). +One model family (gemma-4), two scales, three task families. Hard buckets +are small (n=55/38/25); within-ladder orderings are not individually +significant, and only the pooled hard effect and the HumanEval overall gain +survive multiple-comparison scrutiny. Bucket membership derives from greedy +labeling runs (consensus-k0 robustness check moves numbers <2 points, but +both checks share the base model). Best-of-3/budget-CoT lack per-item logs +(no paired tests against them). The L23 exit arm and a third architecture +family were not run; LiveCodeBench (contamination-safe) was not run; rung-2 +was not run at 12B. The easy-item perturbation tax persists wherever the +gate's precision fails. MBPP/GSM8K likely overlap pretraining data; both +arms share contamination, and memorized items land in the easy bucket, but +bucket composition is contamination-sensitive. The capability panel +(§3.8) is pending; until it lands, off-task effects of k>0 are unmeasured. +The Blocksworld-12B and GSM8K-12B failures mean the adaptive-α fix is +demonstrated on one task at one scale, not established as a general recipe. -## 6. Outlook +## 6. Conclusion -The retrofit recipe — lens-locate, anchor-merge, verifier-filtered -curriculum, gate — is scale-portable by construction: trainable mass is -independent of base size, and prompt-side loops are prefill-shaped, so their -economics *improve* with scale while serial CoT decode gets slower. The open -question that decides whether this is a curiosity or a method is whether the -effect survives scale (12B next; then a mid-size uptraining of the band -itself). If it does, "loopification" becomes a cheap post-training phase any -holder of a pretrained model can apply — a silent planning mode for the -installed base, with its latent reasoning legible to the same lens that -built it. +The experiment this program set out to run — *can an interpretability lens +tell you where to install recurrence in a frozen model, and does it work?* — +has a clean answer: yes, and the placement is causally load-bearing. The +more interesting answer is what the recurrence turned out to be: not a +reasoning engine, but a remarkably cheap way to make a frozen model amortize +its own planning into 0.03% of extra parameters, with a training-time loop +as scaffold and an inference-time loop that is optional once the content +exists. The practical recipe that survives all controls: lens-locate the +band; anchor-merge with a state-dependent coefficient; label difficulty by +STaR; distill plans if you have them, loop if you don't; gate by predicted +difficulty; keep k=0 as the exact base model. What it buys: the +plan-dependent slice at zero tokens and zero decode cost. What it does not +buy: overall accuracy beyond what matched-compute sampling already delivers. +Both halves of that sentence are the contribution. diff --git a/results-loop/STATS.md b/results-loop/STATS.md new file mode 100644 index 0000000..422e2cb --- /dev/null +++ b/results-loop/STATS.md @@ -0,0 +1,47 @@ +# Final statistics pass + +## Headline numbers (Wilson 95% CIs) + +- **MBPP loop s0 k=0 (base)**: overall 0.518 [0.474, 0.561] (n=500); hard 0.055 [0.019, 0.149] (n=55) +- **MBPP loop s0 k=2**: overall 0.530 [0.486, 0.573] (n=500); hard 0.309 [0.203, 0.440] (n=55) +- **MBPP loop s0 k=4**: overall 0.536 [0.492, 0.579] (n=500); hard 0.436 [0.314, 0.567] (n=55) +- **MBPP distill s1 k=1 (FF)**: overall 0.544 [0.500, 0.587] (n=500); hard 0.418 [0.297, 0.550] (n=55) +- **MBPP pause16 k=1**: overall 0.552 [0.508, 0.595] (n=500); hard 0.364 [0.249, 0.496] (n=55) +- **MBPP stack-train k=4**: overall 0.524 [0.480, 0.567] (n=500); hard 0.345 [0.234, 0.477] (n=55) +- **MBPP distill-in-loopmode k=2**: overall 0.458 [0.415, 0.502] (n=500); hard 0.200 [0.116, 0.324] (n=55) +- **Rust transfer k=0**: overall 0.591 [0.512, 0.665] (n=154); hard 0.080 [0.022, 0.250] (n=25) +- **Rust transfer k=4**: overall 0.552 [0.473, 0.628] (n=154); hard 0.240 [0.115, 0.434] (n=25) +- **HumanEval loop k=4**: overall 0.665 [0.589, 0.732] (n=164); hard 0.316 [0.191, 0.475] (n=38) +- **HumanEval distill k=1**: overall 0.646 [0.571, 0.715] (n=164); hard 0.237 [0.130, 0.392] (n=38) +- **MBPP best-of-3 (compute-matched)**: overall 0.572 [0.528, 0.615] (n=500); hard 0.327 [0.218, 0.459] (n=55) +- **MBPP budget-CoT-50**: overall 0.538 [0.494, 0.581] (n=500); hard 0.400 [0.281, 0.532] (n=55) +- **MBPP distill, 8 runs (hard)**: mean 0.457 ± 0.046 sd (range 0.400-0.545); overall mean 0.555 +- **MBPP loop seeds k=4 (hard)**: mean 0.375 ± 0.055 sd (n_seeds=5) + +## McNemar exact tests (paired on items) + +- loop k=4 vs k=0, overall: A-only 30, B-only 39, n=500, p=0.3356 (n.s.) +- loop k=4 vs k=0, hard: A-only 1, B-only 22, n=55, p=5.722e-06 (**significant**) +- distill k=1 vs loop k=4, overall: A-only 33, B-only 37, n=500, p=0.7202 (n.s.) +- distill k=1 vs loop k=4, hard: A-only 10, B-only 9, n=55, p=1 (n.s.) +- stack-train k=4 vs distill k=1, hard: A-only 10, B-only 6, n=55, p=0.4545 (n.s.) +- HumanEval loop k=4 vs k=0, overall: A-only 5, B-only 18, n=164, p=0.01062 (**significant**) +- HumanEval distill k=1 vs k=0, overall: A-only 7, B-only 17, n=164, p=0.06391 (n.s.) +- Rust loop k=4 vs k=0, overall: A-only 11, B-only 5, n=154, p=0.2101 (n.s.) +- Rust loop k=4 vs k=0, hard: A-only 0, B-only 4, n=25, p=0.125 (n.s.) + +## Pooled hard bucket (MBPP + HumanEval + Rust) + +Paired within-item k>0 vs k=0, counts pooled across benchmarks (loop arm; distill pooled where available). + +- **loop**: base 5/118 -> loop 42/118 (0.042 -> 0.356, CI [0.275, 0.446]), McNemar p=1.46e-10 +- **distill**: base 3/93 -> distill 32/93 (0.032 -> 0.344, CI [0.255, 0.445]), McNemar p=2.98e-08 + +## Label robustness (consensus-k0 hard set) + +Hard bucket redefined as: labeled hard AND k=0 fails in every seed's own eval run (removes single-greedy-run selection noise). + +- consensus hard set: 52 of 55 labeled-hard items +- MBPP loop s0 k=4: labeled-hard 0.436 -> consensus-hard 0.423 [0.299, 0.558] (n=52) +- MBPP distill s1 k=1 (FF): labeled-hard 0.418 -> consensus-hard 0.404 [0.282, 0.539] (n=52) +- MBPP stack-train k=4: labeled-hard 0.345 -> consensus-hard 0.346 [0.232, 0.482] (n=52) diff --git a/results-loop/eval_code_distill_s7.json b/results-loop/eval_code_distill_s7.json new file mode 100644 index 0000000..8624716 --- /dev/null +++ b/results-loop/eval_code_distill_s7.json @@ -0,0 +1,4026 @@ +{ + "tag": "distill_s7", + "ks": { + "0": { + "acc": 0.518, + "by_label": { + "easy": 0.9769230769230769, + "hard": 0.05454545454545454, + "drop": 0.010810810810810811 + }, + "per_item": [ + { + "task_id": 11, + "ok": true + }, + { + "task_id": 12, + "ok": false + }, + { + "task_id": 13, + "ok": false + }, + { + "task_id": 14, + "ok": true + }, + { + "task_id": 15, + "ok": false + }, + { + "task_id": 16, + "ok": false + }, + { + "task_id": 17, + "ok": true + }, + { + "task_id": 18, + "ok": true + }, + { + "task_id": 19, + "ok": true + }, + { + "task_id": 20, + "ok": false + }, + { + "task_id": 21, + "ok": false + }, + { + "task_id": 22, + "ok": true + }, + { + "task_id": 23, + "ok": true + }, + { + "task_id": 24, + "ok": false + }, + { + "task_id": 25, + "ok": true + }, + { + "task_id": 26, + "ok": false + }, + { + "task_id": 27, + "ok": true + }, + { + "task_id": 28, + "ok": true + }, + { + "task_id": 29, + "ok": true + }, + { + "task_id": 30, + "ok": true + }, + { + "task_id": 31, + "ok": false + }, + { + "task_id": 32, + "ok": true + }, + { + "task_id": 33, + "ok": false + }, + { + "task_id": 34, + "ok": false + }, + { + "task_id": 35, + "ok": false + }, + { + "task_id": 36, + "ok": false + }, + { + "task_id": 37, + "ok": true + }, + { + "task_id": 38, + "ok": false + }, + { + "task_id": 39, + "ok": false + }, + { + "task_id": 40, + "ok": true + }, + { + "task_id": 41, + "ok": true + }, + { + "task_id": 42, + "ok": false + }, + { + "task_id": 43, + "ok": false + }, + { + "task_id": 44, + "ok": true + }, + { + "task_id": 45, + "ok": true + }, + { + "task_id": 46, + "ok": true + }, + { + "task_id": 47, + "ok": false + }, + { + "task_id": 48, + "ok": false + }, + { + "task_id": 49, + "ok": true + }, + { + "task_id": 50, + "ok": true + }, + { + "task_id": 51, + "ok": true + }, + { + "task_id": 52, + "ok": true + }, + { + "task_id": 53, + "ok": true + }, + { + "task_id": 54, + "ok": false + }, + { + "task_id": 55, + "ok": true + }, + { + "task_id": 56, + "ok": true + }, + { + "task_id": 57, + "ok": true + }, + { + "task_id": 58, + "ok": true + }, + { + "task_id": 59, + "ok": false + }, + { + "task_id": 60, + "ok": false + }, + { + "task_id": 61, + "ok": true + }, + { + "task_id": 62, + "ok": true + }, + { + "task_id": 63, + "ok": true + }, + { + "task_id": 64, + "ok": true + }, + { + "task_id": 65, + "ok": true + }, + { + "task_id": 66, + "ok": true + }, + { + "task_id": 67, + "ok": false + }, + { + "task_id": 68, + "ok": true + }, + { + "task_id": 69, + "ok": true + }, + { + "task_id": 70, + "ok": true + }, + { + "task_id": 71, + "ok": false + }, + { + "task_id": 72, + "ok": false + }, + { + "task_id": 73, + "ok": false + }, + { + "task_id": 74, + "ok": false + }, + { + "task_id": 75, + "ok": true + }, + { + "task_id": 76, + "ok": true + }, + { + "task_id": 77, + "ok": false + }, + { + "task_id": 78, + "ok": true + }, + { + "task_id": 79, + "ok": true + }, + { + "task_id": 80, + "ok": true + }, + { + "task_id": 81, + "ok": false + }, + { + "task_id": 82, + "ok": true + }, + { + "task_id": 83, + "ok": false + }, + { + "task_id": 84, + "ok": false + }, + { + "task_id": 85, + "ok": true + }, + { + "task_id": 86, + "ok": false + }, + { + "task_id": 87, + "ok": false + }, + { + "task_id": 88, + "ok": true + }, + { + "task_id": 89, + "ok": false + }, + { + "task_id": 90, + "ok": true + }, + { + "task_id": 91, + "ok": true + }, + { + "task_id": 92, + "ok": false + }, + { + "task_id": 93, + "ok": true + }, + { + "task_id": 94, + "ok": true + }, + { + "task_id": 95, + "ok": false + }, + { + "task_id": 96, + "ok": true + }, + { + "task_id": 97, + "ok": true + }, + { + "task_id": 98, + "ok": true + }, + { + "task_id": 99, + "ok": true + }, + { + "task_id": 100, + "ok": false + }, + { + "task_id": 101, + "ok": false + }, + { + "task_id": 102, + "ok": true + }, + { + "task_id": 103, + "ok": false + }, + { + "task_id": 104, + "ok": true + }, + { + "task_id": 105, + "ok": true + }, + { + "task_id": 106, + "ok": false + }, + { + "task_id": 107, + "ok": false + }, + { + "task_id": 108, + "ok": false + }, + { + "task_id": 109, + "ok": false + }, + { + "task_id": 110, + "ok": false + }, + { + "task_id": 111, + "ok": false + }, + { + "task_id": 112, + "ok": false + }, + { + "task_id": 113, + "ok": true + }, + { + "task_id": 114, + "ok": false + }, + { + "task_id": 115, + "ok": true + }, + { + "task_id": 116, + "ok": true + }, + { + "task_id": 117, + "ok": true + }, + { + "task_id": 118, + "ok": false + }, + { + "task_id": 119, + "ok": false + }, + { + "task_id": 120, + "ok": false + }, + { + "task_id": 121, + "ok": false + }, + { + "task_id": 122, + "ok": false + }, + { + "task_id": 123, + "ok": false + }, + { + "task_id": 124, + "ok": false + }, + { + "task_id": 125, + "ok": false + }, + { + "task_id": 126, + "ok": false + }, + { + "task_id": 127, + "ok": true + }, + { + "task_id": 128, + "ok": true + }, + { + "task_id": 129, + "ok": false + }, + { + "task_id": 130, + "ok": false + }, + { + "task_id": 131, + "ok": true + }, + { + "task_id": 132, + "ok": true + }, + { + "task_id": 133, + "ok": true + }, + { + "task_id": 134, + "ok": false + }, + { + "task_id": 135, + "ok": true + }, + { + "task_id": 136, + "ok": false + }, + { + "task_id": 137, + "ok": false + }, + { + "task_id": 138, + "ok": false + }, + { + "task_id": 139, + "ok": false + }, + { + "task_id": 140, + "ok": false + }, + { + "task_id": 141, + "ok": false + }, + { + "task_id": 142, + "ok": false + }, + { + "task_id": 143, + "ok": false + }, + { + "task_id": 144, + "ok": true + }, + { + "task_id": 145, + "ok": true + }, + { + "task_id": 146, + "ok": false + }, + { + "task_id": 147, + "ok": false + }, + { + "task_id": 148, + "ok": false + }, + { + "task_id": 149, + "ok": false + }, + { + "task_id": 150, + "ok": false + }, + { + "task_id": 151, + "ok": true + }, + { + "task_id": 152, + "ok": false + }, + { + "task_id": 153, + "ok": true + }, + { + "task_id": 154, + "ok": true + }, + { + "task_id": 155, + "ok": false + }, + { + "task_id": 156, + "ok": true + }, + { + "task_id": 157, + "ok": true + }, + { + "task_id": 158, + "ok": false + }, + { + "task_id": 159, + "ok": false + }, + { + "task_id": 160, + "ok": false + }, + { + "task_id": 161, + "ok": true + }, + { + "task_id": 162, + "ok": true + }, + { + "task_id": 163, + "ok": true + }, + { + "task_id": 164, + "ok": false + }, + { + "task_id": 165, + "ok": false + }, + { + "task_id": 166, + "ok": false + }, + { + "task_id": 167, + "ok": true + }, + { + "task_id": 168, + "ok": true + }, + { + "task_id": 169, + "ok": false + }, + { + "task_id": 170, + "ok": false + }, + { + "task_id": 171, + "ok": true + }, + { + "task_id": 172, + "ok": true + }, + { + "task_id": 173, + "ok": true + }, + { + "task_id": 174, + "ok": true + }, + { + "task_id": 175, + "ok": true + }, + { + "task_id": 176, + "ok": true + }, + { + "task_id": 177, + "ok": false + }, + { + "task_id": 178, + "ok": true + }, + { + "task_id": 179, + "ok": true + }, + { + "task_id": 180, + "ok": false + }, + { + "task_id": 181, + "ok": false + }, + { + "task_id": 182, + "ok": false + }, + { + "task_id": 183, + "ok": false + }, + { + "task_id": 184, + "ok": false + }, + { + "task_id": 185, + "ok": false + }, + { + "task_id": 186, + "ok": true + }, + { + "task_id": 187, + "ok": false + }, + { + "task_id": 188, + "ok": false + }, + { + "task_id": 189, + "ok": false + }, + { + "task_id": 190, + "ok": false + }, + { + "task_id": 191, + "ok": true + }, + { + "task_id": 192, + "ok": true + }, + { + "task_id": 193, + "ok": false + }, + { + "task_id": 194, + "ok": true + }, + { + "task_id": 195, + "ok": false + }, + { + "task_id": 196, + "ok": true + }, + { + "task_id": 197, + "ok": true + }, + { + "task_id": 198, + "ok": false + }, + { + "task_id": 199, + "ok": true + }, + { + "task_id": 200, + "ok": true + }, + { + "task_id": 201, + "ok": true + }, + { + "task_id": 202, + "ok": true + }, + { + "task_id": 203, + "ok": true + }, + { + "task_id": 204, + "ok": true + }, + { + "task_id": 205, + "ok": false + }, + { + "task_id": 206, + "ok": false + }, + { + "task_id": 207, + "ok": false + }, + { + "task_id": 208, + "ok": false + }, + { + "task_id": 209, + "ok": false + }, + { + "task_id": 210, + "ok": true + }, + { + "task_id": 211, + "ok": false + }, + { + "task_id": 212, + "ok": false + }, + { + "task_id": 213, + "ok": true + }, + { + "task_id": 214, + "ok": true + }, + { + "task_id": 215, + "ok": true + }, + { + "task_id": 216, + "ok": false + }, + { + "task_id": 217, + "ok": true + }, + { + "task_id": 218, + "ok": false + }, + { + "task_id": 219, + "ok": false + }, + { + "task_id": 220, + "ok": false + }, + { + "task_id": 221, + "ok": true + }, + { + "task_id": 222, + "ok": true + }, + { + "task_id": 223, + "ok": false + }, + { + "task_id": 224, + "ok": true + }, + { + "task_id": 225, + "ok": false + }, + { + "task_id": 226, + "ok": true + }, + { + "task_id": 227, + "ok": true + }, + { + "task_id": 228, + "ok": false + }, + { + "task_id": 229, + "ok": false + }, + { + "task_id": 230, + "ok": true + }, + { + "task_id": 231, + "ok": false + }, + { + "task_id": 232, + "ok": true + }, + { + "task_id": 233, + "ok": false + }, + { + "task_id": 234, + "ok": true + }, + { + "task_id": 235, + "ok": false + }, + { + "task_id": 236, + "ok": false + }, + { + "task_id": 237, + "ok": false + }, + { + "task_id": 238, + "ok": true + }, + { + "task_id": 239, + "ok": false + }, + { + "task_id": 240, + "ok": true + }, + { + "task_id": 241, + "ok": false + }, + { + "task_id": 242, + "ok": true + }, + { + "task_id": 243, + "ok": false + }, + { + "task_id": 244, + "ok": true + }, + { + "task_id": 245, + "ok": false + }, + { + "task_id": 246, + "ok": true + }, + { + "task_id": 247, + "ok": false + }, + { + "task_id": 248, + "ok": false + }, + { + "task_id": 249, + "ok": false + }, + { + "task_id": 250, + "ok": true + }, + { + "task_id": 251, + "ok": true + }, + { + "task_id": 252, + "ok": true + }, + { + "task_id": 253, + "ok": true + }, + { + "task_id": 254, + "ok": true + }, + { + "task_id": 255, + "ok": false + }, + { + "task_id": 256, + "ok": false + }, + { + "task_id": 257, + "ok": true + }, + { + "task_id": 258, + "ok": false + }, + { + "task_id": 259, + "ok": false + }, + { + "task_id": 260, + "ok": false + }, + { + "task_id": 261, + "ok": true + }, + { + "task_id": 262, + "ok": true + }, + { + "task_id": 263, + "ok": true + }, + { + "task_id": 264, + "ok": false + }, + { + "task_id": 265, + "ok": true + }, + { + "task_id": 266, + "ok": true + }, + { + "task_id": 267, + "ok": true + }, + { + "task_id": 268, + "ok": false + }, + { + "task_id": 269, + "ok": true + }, + { + "task_id": 270, + "ok": false + }, + { + "task_id": 271, + "ok": true + }, + { + "task_id": 272, + "ok": true + }, + { + "task_id": 273, + "ok": true + }, + { + "task_id": 274, + "ok": true + }, + { + "task_id": 275, + "ok": false + }, + { + "task_id": 276, + "ok": false + }, + { + "task_id": 277, + "ok": false + }, + { + "task_id": 278, + "ok": true + }, + { + "task_id": 279, + "ok": false + }, + { + "task_id": 280, + "ok": true + }, + { + "task_id": 281, + "ok": true + }, + { + "task_id": 282, + "ok": false + }, + { + "task_id": 283, + "ok": true + }, + { + "task_id": 284, + "ok": true + }, + { + "task_id": 285, + "ok": true + }, + { + "task_id": 286, + "ok": false + }, + { + "task_id": 287, + "ok": true + }, + { + "task_id": 288, + "ok": false + }, + { + "task_id": 289, + "ok": false + }, + { + "task_id": 290, + "ok": true + }, + { + "task_id": 291, + "ok": false + }, + { + "task_id": 292, + "ok": true + }, + { + "task_id": 293, + "ok": true + }, + { + "task_id": 294, + "ok": true + }, + { + "task_id": 295, + "ok": false + }, + { + "task_id": 296, + "ok": true + }, + { + "task_id": 297, + "ok": true + }, + { + "task_id": 298, + "ok": false + }, + { + "task_id": 299, + "ok": false + }, + { + "task_id": 300, + "ok": false + }, + { + "task_id": 301, + "ok": true + }, + { + "task_id": 302, + "ok": false + }, + { + "task_id": 303, + "ok": false + }, + { + "task_id": 304, + "ok": false + }, + { + "task_id": 305, + "ok": false + }, + { + "task_id": 306, + "ok": false + }, + { + "task_id": 307, + "ok": false + }, + { + "task_id": 308, + "ok": true + }, + { + "task_id": 309, + "ok": true + }, + { + "task_id": 310, + "ok": false + }, + { + "task_id": 311, + "ok": false + }, + { + "task_id": 312, + "ok": false + }, + { + "task_id": 313, + "ok": false + }, + { + "task_id": 314, + "ok": false + }, + { + "task_id": 315, + "ok": true + }, + { + "task_id": 316, + "ok": true + }, + { + "task_id": 317, + "ok": true + }, + { + "task_id": 318, + "ok": false + }, + { + "task_id": 319, + "ok": false + }, + { + "task_id": 320, + "ok": false + }, + { + "task_id": 321, + "ok": false + }, + { + "task_id": 322, + "ok": true + }, + { + "task_id": 323, + "ok": false + }, + { + "task_id": 324, + "ok": false + }, + { + "task_id": 325, + "ok": false + }, + { + "task_id": 326, + "ok": false + }, + { + "task_id": 327, + "ok": true + }, + { + "task_id": 328, + "ok": false + }, + { + "task_id": 329, + "ok": true + }, + { + "task_id": 330, + "ok": true + }, + { + "task_id": 331, + "ok": false + }, + { + "task_id": 332, + "ok": true + }, + { + "task_id": 333, + "ok": true + }, + { + "task_id": 334, + "ok": true + }, + { + "task_id": 335, + "ok": false + }, + { + "task_id": 336, + "ok": true + }, + { + "task_id": 337, + "ok": false + }, + { + "task_id": 338, + "ok": true + }, + { + "task_id": 339, + "ok": false + }, + { + "task_id": 340, + "ok": true + }, + { + "task_id": 341, + "ok": true + }, + { + "task_id": 342, + "ok": false + }, + { + "task_id": 343, + "ok": true + }, + { + "task_id": 344, + "ok": true + }, + { + "task_id": 345, + "ok": true + }, + { + "task_id": 346, + "ok": false + }, + { + "task_id": 347, + "ok": true + }, + { + "task_id": 348, + "ok": false + }, + { + "task_id": 349, + "ok": true + }, + { + "task_id": 350, + "ok": true + }, + { + "task_id": 351, + "ok": false + }, + { + "task_id": 352, + "ok": true + }, + { + "task_id": 353, + "ok": true + }, + { + "task_id": 354, + "ok": true + }, + { + "task_id": 355, + "ok": false + }, + { + "task_id": 356, + "ok": true + }, + { + "task_id": 357, + "ok": true + }, + { + "task_id": 358, + "ok": true + }, + { + "task_id": 359, + "ok": false + }, + { + "task_id": 360, + "ok": false + }, + { + "task_id": 361, + "ok": false + }, + { + "task_id": 362, + "ok": false + }, + { + "task_id": 363, + "ok": true + }, + { + "task_id": 364, + "ok": false + }, + { + "task_id": 365, + "ok": true + }, + { + "task_id": 366, + "ok": true + }, + { + "task_id": 367, + "ok": false + }, + { + "task_id": 368, + "ok": true + }, + { + "task_id": 369, + "ok": false + }, + { + "task_id": 370, + "ok": false + }, + { + "task_id": 371, + "ok": false + }, + { + "task_id": 372, + "ok": true + }, + { + "task_id": 373, + "ok": true + }, + { + "task_id": 374, + "ok": false + }, + { + "task_id": 375, + "ok": false + }, + { + "task_id": 376, + "ok": false + }, + { + "task_id": 377, + "ok": true + }, + { + "task_id": 378, + "ok": true + }, + { + "task_id": 379, + "ok": true + }, + { + "task_id": 380, + "ok": false + }, + { + "task_id": 381, + "ok": true + }, + { + "task_id": 382, + "ok": false + }, + { + "task_id": 383, + "ok": false + }, + { + "task_id": 384, + "ok": false + }, + { + "task_id": 385, + "ok": true + }, + { + "task_id": 386, + "ok": false + }, + { + "task_id": 387, + "ok": true + }, + { + "task_id": 388, + "ok": true + }, + { + "task_id": 389, + "ok": true + }, + { + "task_id": 390, + "ok": false + }, + { + "task_id": 391, + "ok": true + }, + { + "task_id": 392, + "ok": false + }, + { + "task_id": 393, + "ok": false + }, + { + "task_id": 394, + "ok": true + }, + { + "task_id": 395, + "ok": true + }, + { + "task_id": 396, + "ok": false + }, + { + "task_id": 397, + "ok": true + }, + { + "task_id": 398, + "ok": false + }, + { + "task_id": 399, + "ok": true + }, + { + "task_id": 400, + "ok": false + }, + { + "task_id": 401, + "ok": true + }, + { + "task_id": 402, + "ok": false + }, + { + "task_id": 403, + "ok": true + }, + { + "task_id": 404, + "ok": true + }, + { + "task_id": 405, + "ok": true + }, + { + "task_id": 406, + "ok": true + }, + { + "task_id": 407, + "ok": false + }, + { + "task_id": 408, + "ok": false + }, + { + "task_id": 409, + "ok": true + }, + { + "task_id": 410, + "ok": false + }, + { + "task_id": 411, + "ok": true + }, + { + "task_id": 412, + "ok": true + }, + { + "task_id": 413, + "ok": true + }, + { + "task_id": 414, + "ok": true + }, + { + "task_id": 415, + "ok": false + }, + { + "task_id": 416, + "ok": false + }, + { + "task_id": 417, + "ok": false + }, + { + "task_id": 418, + "ok": true + }, + { + "task_id": 419, + "ok": true + }, + { + "task_id": 420, + "ok": true + }, + { + "task_id": 421, + "ok": true + }, + { + "task_id": 422, + "ok": true + }, + { + "task_id": 423, + "ok": false + }, + { + "task_id": 424, + "ok": true + }, + { + "task_id": 425, + "ok": true + }, + { + "task_id": 426, + "ok": true + }, + { + "task_id": 427, + "ok": true + }, + { + "task_id": 428, + "ok": true + }, + { + "task_id": 429, + "ok": false + }, + { + "task_id": 430, + "ok": false + }, + { + "task_id": 431, + "ok": true + }, + { + "task_id": 432, + "ok": false + }, + { + "task_id": 433, + "ok": true + }, + { + "task_id": 434, + "ok": true + }, + { + "task_id": 435, + "ok": true + }, + { + "task_id": 436, + "ok": false + }, + { + "task_id": 437, + "ok": false + }, + { + "task_id": 438, + "ok": false + }, + { + "task_id": 439, + "ok": true + }, + { + "task_id": 440, + "ok": false + }, + { + "task_id": 441, + "ok": true + }, + { + "task_id": 442, + "ok": false + }, + { + "task_id": 443, + "ok": false + }, + { + "task_id": 444, + "ok": false + }, + { + "task_id": 445, + "ok": true + }, + { + "task_id": 446, + "ok": true + }, + { + "task_id": 447, + "ok": true + }, + { + "task_id": 448, + "ok": false + }, + { + "task_id": 449, + "ok": false + }, + { + "task_id": 450, + "ok": true + }, + { + "task_id": 451, + "ok": true + }, + { + "task_id": 452, + "ok": false + }, + { + "task_id": 453, + "ok": true + }, + { + "task_id": 454, + "ok": true + }, + { + "task_id": 455, + "ok": true + }, + { + "task_id": 456, + "ok": true + }, + { + "task_id": 457, + "ok": true + }, + { + "task_id": 458, + "ok": true + }, + { + "task_id": 459, + "ok": true + }, + { + "task_id": 460, + "ok": true + }, + { + "task_id": 461, + "ok": false + }, + { + "task_id": 462, + "ok": false + }, + { + "task_id": 463, + "ok": false + }, + { + "task_id": 464, + "ok": true + }, + { + "task_id": 465, + "ok": true + }, + { + "task_id": 466, + "ok": true + }, + { + "task_id": 467, + "ok": false + }, + { + "task_id": 468, + "ok": false + }, + { + "task_id": 469, + "ok": false + }, + { + "task_id": 470, + "ok": true + }, + { + "task_id": 471, + "ok": true + }, + { + "task_id": 472, + "ok": false + }, + { + "task_id": 473, + "ok": false + }, + { + "task_id": 474, + "ok": true + }, + { + "task_id": 475, + "ok": true + }, + { + "task_id": 476, + "ok": true + }, + { + "task_id": 477, + "ok": true + }, + { + "task_id": 478, + "ok": false + }, + { + "task_id": 479, + "ok": true + }, + { + "task_id": 480, + "ok": false + }, + { + "task_id": 481, + "ok": false + }, + { + "task_id": 482, + "ok": true + }, + { + "task_id": 483, + "ok": false + }, + { + "task_id": 484, + "ok": true + }, + { + "task_id": 485, + "ok": true + }, + { + "task_id": 486, + "ok": false + }, + { + "task_id": 487, + "ok": true + }, + { + "task_id": 488, + "ok": false + }, + { + "task_id": 489, + "ok": false + }, + { + "task_id": 490, + "ok": false + }, + { + "task_id": 491, + "ok": true + }, + { + "task_id": 492, + "ok": true + }, + { + "task_id": 493, + "ok": false + }, + { + "task_id": 494, + "ok": false + }, + { + "task_id": 495, + "ok": true + }, + { + "task_id": 496, + "ok": false + }, + { + "task_id": 497, + "ok": false + }, + { + "task_id": 498, + "ok": true + }, + { + "task_id": 499, + "ok": true + }, + { + "task_id": 500, + "ok": false + }, + { + "task_id": 501, + "ok": false + }, + { + "task_id": 502, + "ok": true + }, + { + "task_id": 503, + "ok": false + }, + { + "task_id": 504, + "ok": true + }, + { + "task_id": 505, + "ok": true + }, + { + "task_id": 506, + "ok": true + }, + { + "task_id": 507, + "ok": true + }, + { + "task_id": 508, + "ok": false + }, + { + "task_id": 509, + "ok": true + }, + { + "task_id": 510, + "ok": true + } + ] + }, + "1": { + "acc": 0.55, + "by_label": { + "easy": 0.9153846153846154, + "hard": 0.4, + "drop": 0.08108108108108109 + }, + "per_item": [ + { + "task_id": 11, + "ok": true + }, + { + "task_id": 12, + "ok": true + }, + { + "task_id": 13, + "ok": false + }, + { + "task_id": 14, + "ok": true + }, + { + "task_id": 15, + "ok": false + }, + { + "task_id": 16, + "ok": false + }, + { + "task_id": 17, + "ok": true + }, + { + "task_id": 18, + "ok": true + }, + { + "task_id": 19, + "ok": true + }, + { + "task_id": 20, + "ok": false + }, + { + "task_id": 21, + "ok": false + }, + { + "task_id": 22, + "ok": true + }, + { + "task_id": 23, + "ok": true + }, + { + "task_id": 24, + "ok": false + }, + { + "task_id": 25, + "ok": true + }, + { + "task_id": 26, + "ok": false + }, + { + "task_id": 27, + "ok": true + }, + { + "task_id": 28, + "ok": true + }, + { + "task_id": 29, + "ok": true + }, + { + "task_id": 30, + "ok": true + }, + { + "task_id": 31, + "ok": false + }, + { + "task_id": 32, + "ok": true + }, + { + "task_id": 33, + "ok": false + }, + { + "task_id": 34, + "ok": false + }, + { + "task_id": 35, + "ok": true + }, + { + "task_id": 36, + "ok": false + }, + { + "task_id": 37, + "ok": true + }, + { + "task_id": 38, + "ok": false + }, + { + "task_id": 39, + "ok": false + }, + { + "task_id": 40, + "ok": true + }, + { + "task_id": 41, + "ok": true + }, + { + "task_id": 42, + "ok": false + }, + { + "task_id": 43, + "ok": false + }, + { + "task_id": 44, + "ok": false + }, + { + "task_id": 45, + "ok": true + }, + { + "task_id": 46, + "ok": true + }, + { + "task_id": 47, + "ok": false + }, + { + "task_id": 48, + "ok": false + }, + { + "task_id": 49, + "ok": true + }, + { + "task_id": 50, + "ok": true + }, + { + "task_id": 51, + "ok": true + }, + { + "task_id": 52, + "ok": true + }, + { + "task_id": 53, + "ok": true + }, + { + "task_id": 54, + "ok": false + }, + { + "task_id": 55, + "ok": true + }, + { + "task_id": 56, + "ok": true + }, + { + "task_id": 57, + "ok": true + }, + { + "task_id": 58, + "ok": true + }, + { + "task_id": 59, + "ok": false + }, + { + "task_id": 60, + "ok": false + }, + { + "task_id": 61, + "ok": true + }, + { + "task_id": 62, + "ok": true + }, + { + "task_id": 63, + "ok": false + }, + { + "task_id": 64, + "ok": true + }, + { + "task_id": 65, + "ok": true + }, + { + "task_id": 66, + "ok": true + }, + { + "task_id": 67, + "ok": false + }, + { + "task_id": 68, + "ok": true + }, + { + "task_id": 69, + "ok": true + }, + { + "task_id": 70, + "ok": true + }, + { + "task_id": 71, + "ok": false + }, + { + "task_id": 72, + "ok": false + }, + { + "task_id": 73, + "ok": false + }, + { + "task_id": 74, + "ok": false + }, + { + "task_id": 75, + "ok": true + }, + { + "task_id": 76, + "ok": true + }, + { + "task_id": 77, + "ok": false + }, + { + "task_id": 78, + "ok": true + }, + { + "task_id": 79, + "ok": true + }, + { + "task_id": 80, + "ok": true + }, + { + "task_id": 81, + "ok": false + }, + { + "task_id": 82, + "ok": true + }, + { + "task_id": 83, + "ok": false + }, + { + "task_id": 84, + "ok": false + }, + { + "task_id": 85, + "ok": true + }, + { + "task_id": 86, + "ok": true + }, + { + "task_id": 87, + "ok": false + }, + { + "task_id": 88, + "ok": true + }, + { + "task_id": 89, + "ok": true + }, + { + "task_id": 90, + "ok": true + }, + { + "task_id": 91, + "ok": true + }, + { + "task_id": 92, + "ok": false + }, + { + "task_id": 93, + "ok": true + }, + { + "task_id": 94, + "ok": true + }, + { + "task_id": 95, + "ok": true + }, + { + "task_id": 96, + "ok": true + }, + { + "task_id": 97, + "ok": true + }, + { + "task_id": 98, + "ok": true + }, + { + "task_id": 99, + "ok": true + }, + { + "task_id": 100, + "ok": false + }, + { + "task_id": 101, + "ok": false + }, + { + "task_id": 102, + "ok": true + }, + { + "task_id": 103, + "ok": false + }, + { + "task_id": 104, + "ok": true + }, + { + "task_id": 105, + "ok": true + }, + { + "task_id": 106, + "ok": false + }, + { + "task_id": 107, + "ok": false + }, + { + "task_id": 108, + "ok": false + }, + { + "task_id": 109, + "ok": false + }, + { + "task_id": 110, + "ok": false + }, + { + "task_id": 111, + "ok": false + }, + { + "task_id": 112, + "ok": false + }, + { + "task_id": 113, + "ok": true + }, + { + "task_id": 114, + "ok": false + }, + { + "task_id": 115, + "ok": true + }, + { + "task_id": 116, + "ok": true + }, + { + "task_id": 117, + "ok": true + }, + { + "task_id": 118, + "ok": true + }, + { + "task_id": 119, + "ok": false + }, + { + "task_id": 120, + "ok": true + }, + { + "task_id": 121, + "ok": false + }, + { + "task_id": 122, + "ok": false + }, + { + "task_id": 123, + "ok": false + }, + { + "task_id": 124, + "ok": false + }, + { + "task_id": 125, + "ok": false + }, + { + "task_id": 126, + "ok": false + }, + { + "task_id": 127, + "ok": true + }, + { + "task_id": 128, + "ok": true + }, + { + "task_id": 129, + "ok": false + }, + { + "task_id": 130, + "ok": false + }, + { + "task_id": 131, + "ok": true + }, + { + "task_id": 132, + "ok": true + }, + { + "task_id": 133, + "ok": true + }, + { + "task_id": 134, + "ok": false + }, + { + "task_id": 135, + "ok": true + }, + { + "task_id": 136, + "ok": false + }, + { + "task_id": 137, + "ok": false + }, + { + "task_id": 138, + "ok": false + }, + { + "task_id": 139, + "ok": false + }, + { + "task_id": 140, + "ok": false + }, + { + "task_id": 141, + "ok": true + }, + { + "task_id": 142, + "ok": false + }, + { + "task_id": 143, + "ok": false + }, + { + "task_id": 144, + "ok": true + }, + { + "task_id": 145, + "ok": true + }, + { + "task_id": 146, + "ok": false + }, + { + "task_id": 147, + "ok": false + }, + { + "task_id": 148, + "ok": false + }, + { + "task_id": 149, + "ok": true + }, + { + "task_id": 150, + "ok": false + }, + { + "task_id": 151, + "ok": true + }, + { + "task_id": 152, + "ok": true + }, + { + "task_id": 153, + "ok": true + }, + { + "task_id": 154, + "ok": true + }, + { + "task_id": 155, + "ok": false + }, + { + "task_id": 156, + "ok": true + }, + { + "task_id": 157, + "ok": true + }, + { + "task_id": 158, + "ok": false + }, + { + "task_id": 159, + "ok": false + }, + { + "task_id": 160, + "ok": false + }, + { + "task_id": 161, + "ok": true + }, + { + "task_id": 162, + "ok": true + }, + { + "task_id": 163, + "ok": true + }, + { + "task_id": 164, + "ok": false + }, + { + "task_id": 165, + "ok": true + }, + { + "task_id": 166, + "ok": true + }, + { + "task_id": 167, + "ok": true + }, + { + "task_id": 168, + "ok": true + }, + { + "task_id": 169, + "ok": false + }, + { + "task_id": 170, + "ok": false + }, + { + "task_id": 171, + "ok": true + }, + { + "task_id": 172, + "ok": true + }, + { + "task_id": 173, + "ok": true + }, + { + "task_id": 174, + "ok": true + }, + { + "task_id": 175, + "ok": true + }, + { + "task_id": 176, + "ok": true + }, + { + "task_id": 177, + "ok": false + }, + { + "task_id": 178, + "ok": true + }, + { + "task_id": 179, + "ok": true + }, + { + "task_id": 180, + "ok": false + }, + { + "task_id": 181, + "ok": false + }, + { + "task_id": 182, + "ok": false + }, + { + "task_id": 183, + "ok": false + }, + { + "task_id": 184, + "ok": false + }, + { + "task_id": 185, + "ok": false + }, + { + "task_id": 186, + "ok": true + }, + { + "task_id": 187, + "ok": false + }, + { + "task_id": 188, + "ok": false + }, + { + "task_id": 189, + "ok": false + }, + { + "task_id": 190, + "ok": false + }, + { + "task_id": 191, + "ok": true + }, + { + "task_id": 192, + "ok": true + }, + { + "task_id": 193, + "ok": true + }, + { + "task_id": 194, + "ok": true + }, + { + "task_id": 195, + "ok": true + }, + { + "task_id": 196, + "ok": true + }, + { + "task_id": 197, + "ok": true + }, + { + "task_id": 198, + "ok": false + }, + { + "task_id": 199, + "ok": true + }, + { + "task_id": 200, + "ok": true + }, + { + "task_id": 201, + "ok": true + }, + { + "task_id": 202, + "ok": true + }, + { + "task_id": 203, + "ok": true + }, + { + "task_id": 204, + "ok": true + }, + { + "task_id": 205, + "ok": false + }, + { + "task_id": 206, + "ok": false + }, + { + "task_id": 207, + "ok": false + }, + { + "task_id": 208, + "ok": false + }, + { + "task_id": 209, + "ok": false + }, + { + "task_id": 210, + "ok": true + }, + { + "task_id": 211, + "ok": false + }, + { + "task_id": 212, + "ok": true + }, + { + "task_id": 213, + "ok": false + }, + { + "task_id": 214, + "ok": true + }, + { + "task_id": 215, + "ok": true + }, + { + "task_id": 216, + "ok": false + }, + { + "task_id": 217, + "ok": true + }, + { + "task_id": 218, + "ok": false + }, + { + "task_id": 219, + "ok": false + }, + { + "task_id": 220, + "ok": false + }, + { + "task_id": 221, + "ok": true + }, + { + "task_id": 222, + "ok": true + }, + { + "task_id": 223, + "ok": false + }, + { + "task_id": 224, + "ok": true + }, + { + "task_id": 225, + "ok": true + }, + { + "task_id": 226, + "ok": true + }, + { + "task_id": 227, + "ok": true + }, + { + "task_id": 228, + "ok": false + }, + { + "task_id": 229, + "ok": false + }, + { + "task_id": 230, + "ok": false + }, + { + "task_id": 231, + "ok": false + }, + { + "task_id": 232, + "ok": true + }, + { + "task_id": 233, + "ok": false + }, + { + "task_id": 234, + "ok": true + }, + { + "task_id": 235, + "ok": false + }, + { + "task_id": 236, + "ok": false + }, + { + "task_id": 237, + "ok": false + }, + { + "task_id": 238, + "ok": true + }, + { + "task_id": 239, + "ok": false + }, + { + "task_id": 240, + "ok": true + }, + { + "task_id": 241, + "ok": false + }, + { + "task_id": 242, + "ok": true + }, + { + "task_id": 243, + "ok": false + }, + { + "task_id": 244, + "ok": true + }, + { + "task_id": 245, + "ok": false + }, + { + "task_id": 246, + "ok": true + }, + { + "task_id": 247, + "ok": true + }, + { + "task_id": 248, + "ok": false + }, + { + "task_id": 249, + "ok": false + }, + { + "task_id": 250, + "ok": true + }, + { + "task_id": 251, + "ok": true + }, + { + "task_id": 252, + "ok": true + }, + { + "task_id": 253, + "ok": true + }, + { + "task_id": 254, + "ok": false + }, + { + "task_id": 255, + "ok": false + }, + { + "task_id": 256, + "ok": true + }, + { + "task_id": 257, + "ok": true + }, + { + "task_id": 258, + "ok": true + }, + { + "task_id": 259, + "ok": false + }, + { + "task_id": 260, + "ok": false + }, + { + "task_id": 261, + "ok": true + }, + { + "task_id": 262, + "ok": true + }, + { + "task_id": 263, + "ok": true + }, + { + "task_id": 264, + "ok": true + }, + { + "task_id": 265, + "ok": false + }, + { + "task_id": 266, + "ok": true + }, + { + "task_id": 267, + "ok": true + }, + { + "task_id": 268, + "ok": false + }, + { + "task_id": 269, + "ok": true + }, + { + "task_id": 270, + "ok": true + }, + { + "task_id": 271, + "ok": true + }, + { + "task_id": 272, + "ok": true + }, + { + "task_id": 273, + "ok": true + }, + { + "task_id": 274, + "ok": true + }, + { + "task_id": 275, + "ok": false + }, + { + "task_id": 276, + "ok": false + }, + { + "task_id": 277, + "ok": false + }, + { + "task_id": 278, + "ok": true + }, + { + "task_id": 279, + "ok": false + }, + { + "task_id": 280, + "ok": true + }, + { + "task_id": 281, + "ok": true + }, + { + "task_id": 282, + "ok": false + }, + { + "task_id": 283, + "ok": true + }, + { + "task_id": 284, + "ok": true + }, + { + "task_id": 285, + "ok": true + }, + { + "task_id": 286, + "ok": false + }, + { + "task_id": 287, + "ok": true + }, + { + "task_id": 288, + "ok": false + }, + { + "task_id": 289, + "ok": false + }, + { + "task_id": 290, + "ok": true + }, + { + "task_id": 291, + "ok": false + }, + { + "task_id": 292, + "ok": true + }, + { + "task_id": 293, + "ok": true + }, + { + "task_id": 294, + "ok": true + }, + { + "task_id": 295, + "ok": false + }, + { + "task_id": 296, + "ok": false + }, + { + "task_id": 297, + "ok": true + }, + { + "task_id": 298, + "ok": false + }, + { + "task_id": 299, + "ok": false + }, + { + "task_id": 300, + "ok": false + }, + { + "task_id": 301, + "ok": false + }, + { + "task_id": 302, + "ok": true + }, + { + "task_id": 303, + "ok": false + }, + { + "task_id": 304, + "ok": false + }, + { + "task_id": 305, + "ok": false + }, + { + "task_id": 306, + "ok": false + }, + { + "task_id": 307, + "ok": false + }, + { + "task_id": 308, + "ok": false + }, + { + "task_id": 309, + "ok": true + }, + { + "task_id": 310, + "ok": false + }, + { + "task_id": 311, + "ok": false + }, + { + "task_id": 312, + "ok": false + }, + { + "task_id": 313, + "ok": false + }, + { + "task_id": 314, + "ok": false + }, + { + "task_id": 315, + "ok": true + }, + { + "task_id": 316, + "ok": true + }, + { + "task_id": 317, + "ok": true + }, + { + "task_id": 318, + "ok": false + }, + { + "task_id": 319, + "ok": true + }, + { + "task_id": 320, + "ok": false + }, + { + "task_id": 321, + "ok": false + }, + { + "task_id": 322, + "ok": true + }, + { + "task_id": 323, + "ok": false + }, + { + "task_id": 324, + "ok": false + }, + { + "task_id": 325, + "ok": false + }, + { + "task_id": 326, + "ok": false + }, + { + "task_id": 327, + "ok": true + }, + { + "task_id": 328, + "ok": false + }, + { + "task_id": 329, + "ok": true + }, + { + "task_id": 330, + "ok": false + }, + { + "task_id": 331, + "ok": false + }, + { + "task_id": 332, + "ok": true + }, + { + "task_id": 333, + "ok": true + }, + { + "task_id": 334, + "ok": true + }, + { + "task_id": 335, + "ok": true + }, + { + "task_id": 336, + "ok": true + }, + { + "task_id": 337, + "ok": false + }, + { + "task_id": 338, + "ok": true + }, + { + "task_id": 339, + "ok": false + }, + { + "task_id": 340, + "ok": true + }, + { + "task_id": 341, + "ok": true + }, + { + "task_id": 342, + "ok": false + }, + { + "task_id": 343, + "ok": false + }, + { + "task_id": 344, + "ok": false + }, + { + "task_id": 345, + "ok": true + }, + { + "task_id": 346, + "ok": false + }, + { + "task_id": 347, + "ok": true + }, + { + "task_id": 348, + "ok": false + }, + { + "task_id": 349, + "ok": true + }, + { + "task_id": 350, + "ok": true + }, + { + "task_id": 351, + "ok": false + }, + { + "task_id": 352, + "ok": true + }, + { + "task_id": 353, + "ok": true + }, + { + "task_id": 354, + "ok": false + }, + { + "task_id": 355, + "ok": false + }, + { + "task_id": 356, + "ok": true + }, + { + "task_id": 357, + "ok": true + }, + { + "task_id": 358, + "ok": true + }, + { + "task_id": 359, + "ok": false + }, + { + "task_id": 360, + "ok": false + }, + { + "task_id": 361, + "ok": true + }, + { + "task_id": 362, + "ok": false + }, + { + "task_id": 363, + "ok": true + }, + { + "task_id": 364, + "ok": false + }, + { + "task_id": 365, + "ok": true + }, + { + "task_id": 366, + "ok": true + }, + { + "task_id": 367, + "ok": false + }, + { + "task_id": 368, + "ok": true + }, + { + "task_id": 369, + "ok": true + }, + { + "task_id": 370, + "ok": false + }, + { + "task_id": 371, + "ok": false + }, + { + "task_id": 372, + "ok": true + }, + { + "task_id": 373, + "ok": true + }, + { + "task_id": 374, + "ok": false + }, + { + "task_id": 375, + "ok": true + }, + { + "task_id": 376, + "ok": false + }, + { + "task_id": 377, + "ok": true + }, + { + "task_id": 378, + "ok": true + }, + { + "task_id": 379, + "ok": true + }, + { + "task_id": 380, + "ok": false + }, + { + "task_id": 381, + "ok": true + }, + { + "task_id": 382, + "ok": false + }, + { + "task_id": 383, + "ok": false + }, + { + "task_id": 384, + "ok": false + }, + { + "task_id": 385, + "ok": false + }, + { + "task_id": 386, + "ok": false + }, + { + "task_id": 387, + "ok": true + }, + { + "task_id": 388, + "ok": true + }, + { + "task_id": 389, + "ok": false + }, + { + "task_id": 390, + "ok": false + }, + { + "task_id": 391, + "ok": true + }, + { + "task_id": 392, + "ok": true + }, + { + "task_id": 393, + "ok": true + }, + { + "task_id": 394, + "ok": true + }, + { + "task_id": 395, + "ok": true + }, + { + "task_id": 396, + "ok": false + }, + { + "task_id": 397, + "ok": true + }, + { + "task_id": 398, + "ok": true + }, + { + "task_id": 399, + "ok": false + }, + { + "task_id": 400, + "ok": false + }, + { + "task_id": 401, + "ok": true + }, + { + "task_id": 402, + "ok": false + }, + { + "task_id": 403, + "ok": true + }, + { + "task_id": 404, + "ok": true + }, + { + "task_id": 405, + "ok": true + }, + { + "task_id": 406, + "ok": true + }, + { + "task_id": 407, + "ok": false + }, + { + "task_id": 408, + "ok": false + }, + { + "task_id": 409, + "ok": true + }, + { + "task_id": 410, + "ok": true + }, + { + "task_id": 411, + "ok": false + }, + { + "task_id": 412, + "ok": true + }, + { + "task_id": 413, + "ok": true + }, + { + "task_id": 414, + "ok": true + }, + { + "task_id": 415, + "ok": false + }, + { + "task_id": 416, + "ok": false + }, + { + "task_id": 417, + "ok": false + }, + { + "task_id": 418, + "ok": true + }, + { + "task_id": 419, + "ok": true + }, + { + "task_id": 420, + "ok": true + }, + { + "task_id": 421, + "ok": true + }, + { + "task_id": 422, + "ok": true + }, + { + "task_id": 423, + "ok": false + }, + { + "task_id": 424, + "ok": true + }, + { + "task_id": 425, + "ok": true + }, + { + "task_id": 426, + "ok": true + }, + { + "task_id": 427, + "ok": true + }, + { + "task_id": 428, + "ok": true + }, + { + "task_id": 429, + "ok": true + }, + { + "task_id": 430, + "ok": false + }, + { + "task_id": 431, + "ok": true + }, + { + "task_id": 432, + "ok": false + }, + { + "task_id": 433, + "ok": true + }, + { + "task_id": 434, + "ok": false + }, + { + "task_id": 435, + "ok": true + }, + { + "task_id": 436, + "ok": false + }, + { + "task_id": 437, + "ok": true + }, + { + "task_id": 438, + "ok": false + }, + { + "task_id": 439, + "ok": false + }, + { + "task_id": 440, + "ok": true + }, + { + "task_id": 441, + "ok": true + }, + { + "task_id": 442, + "ok": false + }, + { + "task_id": 443, + "ok": false + }, + { + "task_id": 444, + "ok": false + }, + { + "task_id": 445, + "ok": true + }, + { + "task_id": 446, + "ok": true + }, + { + "task_id": 447, + "ok": true + }, + { + "task_id": 448, + "ok": false + }, + { + "task_id": 449, + "ok": false + }, + { + "task_id": 450, + "ok": true + }, + { + "task_id": 451, + "ok": true + }, + { + "task_id": 452, + "ok": false + }, + { + "task_id": 453, + "ok": true + }, + { + "task_id": 454, + "ok": true + }, + { + "task_id": 455, + "ok": true + }, + { + "task_id": 456, + "ok": true + }, + { + "task_id": 457, + "ok": true + }, + { + "task_id": 458, + "ok": true + }, + { + "task_id": 459, + "ok": true + }, + { + "task_id": 460, + "ok": true + }, + { + "task_id": 461, + "ok": false + }, + { + "task_id": 462, + "ok": false + }, + { + "task_id": 463, + "ok": false + }, + { + "task_id": 464, + "ok": true + }, + { + "task_id": 465, + "ok": true + }, + { + "task_id": 466, + "ok": true + }, + { + "task_id": 467, + "ok": false + }, + { + "task_id": 468, + "ok": true + }, + { + "task_id": 469, + "ok": false + }, + { + "task_id": 470, + "ok": false + }, + { + "task_id": 471, + "ok": true + }, + { + "task_id": 472, + "ok": false + }, + { + "task_id": 473, + "ok": false + }, + { + "task_id": 474, + "ok": true + }, + { + "task_id": 475, + "ok": true + }, + { + "task_id": 476, + "ok": true + }, + { + "task_id": 477, + "ok": true + }, + { + "task_id": 478, + "ok": false + }, + { + "task_id": 479, + "ok": true + }, + { + "task_id": 480, + "ok": true + }, + { + "task_id": 481, + "ok": false + }, + { + "task_id": 482, + "ok": true + }, + { + "task_id": 483, + "ok": false + }, + { + "task_id": 484, + "ok": true + }, + { + "task_id": 485, + "ok": true + }, + { + "task_id": 486, + "ok": true + }, + { + "task_id": 487, + "ok": true + }, + { + "task_id": 488, + "ok": false + }, + { + "task_id": 489, + "ok": false + }, + { + "task_id": 490, + "ok": false + }, + { + "task_id": 491, + "ok": true + }, + { + "task_id": 492, + "ok": true + }, + { + "task_id": 493, + "ok": false + }, + { + "task_id": 494, + "ok": false + }, + { + "task_id": 495, + "ok": true + }, + { + "task_id": 496, + "ok": false + }, + { + "task_id": 497, + "ok": false + }, + { + "task_id": 498, + "ok": true + }, + { + "task_id": 499, + "ok": true + }, + { + "task_id": 500, + "ok": false + }, + { + "task_id": 501, + "ok": false + }, + { + "task_id": 502, + "ok": true + }, + { + "task_id": 503, + "ok": false + }, + { + "task_id": 504, + "ok": true + }, + { + "task_id": 505, + "ok": true + }, + { + "task_id": 506, + "ok": true + }, + { + "task_id": 507, + "ok": true + }, + { + "task_id": 508, + "ok": false + }, + { + "task_id": 509, + "ok": true + }, + { + "task_id": 510, + "ok": false + } + ] + } + }, + "n": 500 +} \ No newline at end of file diff --git a/results-loop/eval_rung2_s0.json b/results-loop/eval_rung2_s0.json new file mode 100644 index 0000000..08c4410 --- /dev/null +++ b/results-loop/eval_rung2_s0.json @@ -0,0 +1,26 @@ +{ + "0": { + "acc": 0.5, + "by_label": { + "easy": 0.9672131147540983, + "hard": 0.17857142857142858, + "drop": 0.02 + } + }, + "2": { + "acc": 0.504, + "by_label": { + "easy": 0.9098360655737705, + "hard": 0.35714285714285715, + "drop": 0.05 + } + }, + "4": { + "acc": 0.512, + "by_label": { + "easy": 0.9098360655737705, + "hard": 0.42857142857142855, + "drop": 0.05 + } + } +} \ No newline at end of file diff --git a/results-loop/eval_rung2_s1.json b/results-loop/eval_rung2_s1.json new file mode 100644 index 0000000..6a2c365 --- /dev/null +++ b/results-loop/eval_rung2_s1.json @@ -0,0 +1,26 @@ +{ + "0": { + "acc": 0.496, + "by_label": { + "easy": 0.9672131147540983, + "hard": 0.10714285714285714, + "drop": 0.03 + } + }, + "2": { + "acc": 0.516, + "by_label": { + "easy": 0.9344262295081968, + "hard": 0.39285714285714285, + "drop": 0.04 + } + }, + "4": { + "acc": 0.524, + "by_label": { + "easy": 0.9098360655737705, + "hard": 0.4642857142857143, + "drop": 0.07 + } + } +} \ No newline at end of file diff --git a/results-loop/fig_ladder.png b/results-loop/fig_ladder.png new file mode 100644 index 0000000..660a706 Binary files /dev/null and b/results-loop/fig_ladder.png differ diff --git a/results-loop/fig_placement.png b/results-loop/fig_placement.png new file mode 100644 index 0000000..2e1a38d Binary files /dev/null and b/results-loop/fig_placement.png differ diff --git a/results-loop/fig_scale.png b/results-loop/fig_scale.png new file mode 100644 index 0000000..d214c1d Binary files /dev/null and b/results-loop/fig_scale.png differ diff --git a/results-loop/fig_transfer.png b/results-loop/fig_transfer.png new file mode 100644 index 0000000..19e427d Binary files /dev/null and b/results-loop/fig_transfer.png differ diff --git a/results-loop/stats_final.json b/results-loop/stats_final.json new file mode 100644 index 0000000..5278d67 --- /dev/null +++ b/results-loop/stats_final.json @@ -0,0 +1,369 @@ +{ + "MBPP loop s0 k=0 (base)": { + "overall": { + "acc": 0.518, + "n": 500, + "ci": [ + 0.474, + 0.561 + ] + }, + "hard": { + "acc": 0.05454545454545454, + "n": 55, + "ci": [ + 0.019, + 0.149 + ] + } + }, + "MBPP loop s0 k=2": { + "overall": { + "acc": 0.53, + "n": 500, + "ci": [ + 0.486, + 0.573 + ] + }, + "hard": { + "acc": 0.3090909090909091, + "n": 55, + "ci": [ + 0.203, + 0.44 + ] + } + }, + "MBPP loop s0 k=4": { + "overall": { + "acc": 0.536, + "n": 500, + "ci": [ + 0.492, + 0.579 + ] + }, + "hard": { + "acc": 0.43636363636363634, + "n": 55, + "ci": [ + 0.314, + 0.567 + ] + } + }, + "MBPP distill s1 k=1 (FF)": { + "overall": { + "acc": 0.544, + "n": 500, + "ci": [ + 0.5, + 0.587 + ] + }, + "hard": { + "acc": 0.41818181818181815, + "n": 55, + "ci": [ + 0.297, + 0.55 + ] + } + }, + "MBPP pause16 k=1": { + "overall": { + "acc": 0.552, + "n": 500, + "ci": [ + 0.508, + 0.595 + ] + }, + "hard": { + "acc": 0.36363636363636365, + "n": 55, + "ci": [ + 0.249, + 0.496 + ] + } + }, + "MBPP stack-train k=4": { + "overall": { + "acc": 0.524, + "n": 500, + "ci": [ + 0.48, + 0.567 + ] + }, + "hard": { + "acc": 0.34545454545454546, + "n": 55, + "ci": [ + 0.234, + 0.477 + ] + } + }, + "MBPP distill-in-loopmode k=2": { + "overall": { + "acc": 0.458, + "n": 500, + "ci": [ + 0.415, + 0.502 + ] + }, + "hard": { + "acc": 0.2, + "n": 55, + "ci": [ + 0.116, + 0.324 + ] + } + }, + "Rust transfer k=0": { + "overall": { + "acc": 0.5909090909090909, + "n": 154, + "ci": [ + 0.512, + 0.665 + ] + }, + "hard": { + "acc": 0.08, + "n": 25, + "ci": [ + 0.022, + 0.25 + ] + } + }, + "Rust transfer k=4": { + "overall": { + "acc": 0.551948051948052, + "n": 154, + "ci": [ + 0.473, + 0.628 + ] + }, + "hard": { + "acc": 0.24, + "n": 25, + "ci": [ + 0.115, + 0.434 + ] + } + }, + "HumanEval loop k=4": { + "overall": { + "acc": 0.6646341463414634, + "n": 164, + "ci": [ + 0.589, + 0.732 + ] + }, + "hard": { + "acc": 0.3157894736842105, + "n": 38, + "ci": [ + 0.191, + 0.475 + ] + } + }, + "HumanEval distill k=1": { + "overall": { + "acc": 0.6463414634146342, + "n": 164, + "ci": [ + 0.571, + 0.715 + ] + }, + "hard": { + "acc": 0.23684210526315788, + "n": 38, + "ci": [ + 0.13, + 0.392 + ] + } + }, + "MBPP best-of-3 (compute-matched)": { + "overall": { + "acc": 0.572, + "n": 500, + "ci": [ + 0.528, + 0.615 + ] + }, + "hard": { + "acc": 0.32727272727272727, + "n": 55, + "ci": [ + 0.218, + 0.459 + ] + } + }, + "MBPP budget-CoT-50": { + "overall": { + "acc": 0.538, + "n": 500, + "ci": [ + 0.494, + 0.581 + ] + }, + "hard": { + "acc": 0.4, + "n": 55, + "ci": [ + 0.281, + 0.532 + ] + } + }, + "distill_seed_spread": { + "hard": [ + 0.4909090909090909, + 0.41818181818181815, + 0.5454545454545454, + 0.43636363636363634, + 0.45454545454545453, + 0.4727272727272727, + 0.43636363636363634, + 0.4 + ], + "overall": [ + 0.55, + 0.544, + 0.562, + 0.546, + 0.566, + 0.562, + 0.564, + 0.55 + ] + }, + "loop_seed_spread_hard": [ + 0.43636363636363634, + 0.41818181818181815, + 0.3090909090909091, + 0.32727272727272727, + 0.38181818181818183 + ], + "mcnemar": { + "loop k=4 vs k=0, overall": { + "n": 500, + "a_only": 30, + "b_only": 39, + "p": 0.33555761823401514 + }, + "loop k=4 vs k=0, hard": { + "n": 55, + "a_only": 1, + "b_only": 22, + "p": 5.7220458984375e-06 + }, + "distill k=1 vs loop k=4, overall": { + "n": 500, + "a_only": 33, + "b_only": 37, + "p": 0.7202027723528613 + }, + "distill k=1 vs loop k=4, hard": { + "n": 55, + "a_only": 10, + "b_only": 9, + "p": 1.0 + }, + "stack-train k=4 vs distill k=1, hard": { + "n": 55, + "a_only": 10, + "b_only": 6, + "p": 0.454498291015625 + }, + "HumanEval loop k=4 vs k=0, overall": { + "n": 164, + "a_only": 5, + "b_only": 18, + "p": 0.010622024536132812 + }, + "HumanEval distill k=1 vs k=0, overall": { + "n": 164, + "a_only": 7, + "b_only": 17, + "p": 0.06391465663909912 + }, + "Rust loop k=4 vs k=0, overall": { + "n": 154, + "a_only": 11, + "b_only": 5, + "p": 0.210113525390625 + }, + "Rust loop k=4 vs k=0, hard": { + "n": 25, + "a_only": 0, + "b_only": 4, + "p": 0.125 + } + }, + "pooled_hard_loop": { + "base": 5, + "arm": 42, + "n": 118, + "mcnemar": { + "n": 118, + "a_only": 1, + "b_only": 38, + "p": 1.4551915228366852e-10 + } + }, + "pooled_hard_distill": { + "base": 3, + "arm": 32, + "n": 93, + "mcnemar": { + "n": 93, + "a_only": 1, + "b_only": 30, + "p": 2.9802322387695312e-08 + } + }, + "consensus_hard": { + "MBPP loop s0 k=4": { + "acc": 0.4230769230769231, + "n": 52, + "ci": [ + 0.299, + 0.558 + ] + }, + "MBPP distill s1 k=1 (FF)": { + "acc": 0.40384615384615385, + "n": 52, + "ci": [ + 0.282, + 0.539 + ] + }, + "MBPP stack-train k=4": { + "acc": 0.34615384615384615, + "n": 52, + "ci": [ + 0.232, + 0.482 + ] + } + } +} \ No newline at end of file diff --git a/scripts/figures_final.py b/scripts/figures_final.py new file mode 100644 index 0000000..bb48d31 --- /dev/null +++ b/scripts/figures_final.py @@ -0,0 +1,250 @@ +"""Final paper figures (CPU only), regenerated from consolidated results. + + fig_placement.png : band-entrance cliff + structural nulls + fig_ladder.png : MBPP-E2B attribution ladder, hard bucket, Wilson CIs + fig_scale.png : cross-scale / cross-task attribution grid + fig_transfer.png : substrate transfer panel (HumanEval / Rust / BW) +""" + +import json +import math +from pathlib import Path + +import matplotlib +matplotlib.use("Agg") +import matplotlib.pyplot as plt + +ROOT = Path(__file__).resolve().parent.parent +OUT = ROOT / "results-loop" + +BLUE, GREEN, ORANGE, GRAY, RED = ("#2b6cb0", "#2f855a", "#dd6b20", + "#8a8f98", "#c53030") + + +def wilson(c, n, z=1.96): + p = c / n + d = 1 + z * z / n + ctr = (p + z * z / (2 * n)) / d + hw = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d + return ctr - hw, ctr + hw + + +def ev(path, k): + d = json.load(open(path)) + v = d["ks"][str(k)] + return v["acc"], v["by_label"].get("hard", float("nan")) + + +def best_hard(path, exclude0=True): + d = json.load(open(path)) + items = [(int(k), v) for k, v in d["ks"].items() + if not (exclude0 and k == "0")] + k, v = max(items, key=lambda kv: kv[1]["by_label"].get("hard", 0)) + return k, v["acc"], v["by_label"].get("hard", 0) + + +def style(ax): + ax.grid(True, color="#e8e8e8", lw=0.7) + ax.set_axisbelow(True) + for s in ("top", "right"): + ax.spines[s].set_visible(False) + + +# ---------------- fig 1: placement ---------------- +def fig_placement(): + entr = [] # (entrance_layer, overall@bestk, hard@bestk) + for lo in (9, 11, 12, 13, 14): + f = ROOT / f"results-band-{lo}_30" / f"eval_code_band_{lo}_30.json" + if lo == 14: + f = OUT / "eval_code_code_s0_full.json" + k, a, h = best_hard(f) + entr.append((lo, a, h)) + base_a, base_h = ev(OUT / "eval_code_code_s0_full.json", 0) + + fig, ax = plt.subplots(figsize=(6.4, 4.2)) + style(ax) + xs = [e[0] for e in entr] + ax.plot(xs, [e[2] for e in entr], "-s", color=BLUE, lw=2, ms=6, + label="hard bucket (best k)") + ax.plot(xs, [e[1] for e in entr], "-o", color=GRAY, lw=2, ms=5, + label="overall (same k)") + ax.axhline(base_h, color=BLUE, lw=1, ls=":", alpha=0.6) + ax.axhline(base_a, color=GRAY, lw=1, ls=":", alpha=0.6) + ax.annotate("base hard", (9.1, base_h + 0.012), fontsize=8, color=BLUE) + ax.annotate("base overall", (9.1, base_a + 0.012), fontsize=8, color=GRAY) + ax.axvspan(13.5, 14.5, color="#ebf4ff", zorder=0) + ax.annotate("lens-identified\nworkspace entrance", (13.55, 0.60), + fontsize=8, color=BLUE) + ax.set_xticks(xs) + ax.set_xlabel("loop entrance layer (exit fixed at L30)") + ax.set_ylabel("MBPP pass@1") + ax.set_title("Placement cliff: the retrofit works only at the " + "lens boundary (L14)", fontsize=11) + ax.legend(fontsize=8, frameon=False, loc="center left") + fig.text(0.13, 0.005, + "Entrances 17/24 (not shown): structurally null — KV sharing " + "makes k>0 bit-identical to k=0.", fontsize=7.5, color="#666") + fig.tight_layout(rect=(0, 0.03, 1, 1)) + fig.savefig(OUT / "fig_placement.png", dpi=140, facecolor="white", + bbox_inches="tight") + print("wrote fig_placement.png") + + +# ---------------- fig 2: attribution ladder ---------------- +def fig_ladder(): + NH = 55 + rows = [] # (label, hard_acc, n, color, note) + + def add(label, h, n, color, note=""): + rows.append((label, h, n, color, note)) + + _, h0 = ev(OUT / "eval_code_code_s0_full.json", 0) + add("base (k=0, exact)", h0, NH, GRAY) + # untrained E2B control (250-item era, n_hard=28) + d = json.load(open(OUT / "eval_code_untrained.json")) + hu = max(v["by_label"]["hard"] for k, v in d["ks"].items() if k != "0") + add("untrained loop (best k)", hu, 28, GRAY) + d = json.load(open(OUT / "eval_code_ff.json")) + add("trained FF (no recurrence)", d["ks"]["1"]["by_label"]["hard"], 28, + ORANGE) + _, hp = ev(OUT / "eval_code_pause16.json", 1) + add("pause-16 registers (width)", hp, NH, ORANGE) + _, hl = ev(OUT / "eval_code_code_s0_full.json", 4) + add("loop k=4 (depth)", hl, NH, BLUE, "seed mean 0.375 ± 0.055") + d = json.load(open(OUT / "eval_rung2_s1.json")) + add("rung-2: + band LoRA (k=4)", d["4"]["by_label"]["hard"], 28, BLUE) + _, hd = ev(OUT / "eval_code_distill_s1.json", 1) + add("plan-distilled FF", hd, NH, GREEN, "8-run mean 0.457 ± 0.046") + d = json.load(open(OUT / "eval_budgetcot.json")) + add("budget-CoT (50 visible tok)", d["by_label"]["hard"], NH, "#805ad5") + d = json.load(open(OUT / "eval_bestof3.json")) + add("best-of-3 sampling (~matched FLOPs)", d["by_label"]["hard"], NH, + "#805ad5") + d = json.load(open(OUT / "eval_plan_baseline.json")) + add("explicit plan in context (ceiling)", d["by_label"]["hard"], NH, + "#1a202c") + + fig, ax = plt.subplots(figsize=(7.4, 4.8)) + style(ax) + ys = range(len(rows))[::-1] + for y, (label, h, n, color, note) in zip(ys, rows): + lo, hi = wilson(round(h * n), n) + ax.barh(y, h, color=color, height=0.62, alpha=0.88) + ax.plot([lo, hi], [y, y], color="#333", lw=1.2) + txt = f"{h:.2f}" + if note: + txt += f" ({note})" + ax.text(hi + 0.015, y, txt, va="center", fontsize=8) + ax.set_yticks(list(ys)) + ax.set_yticklabels([r[0] for r in rows], fontsize=9) + ax.set_xlim(0, 1.02) + ax.set_xlabel("pass@1, MBPP hard bucket (plan-dependent items)") + ax.set_title("Attribution ladder: what closes the plan gap " + "(bars: point estimate, whiskers: Wilson 95%)", fontsize=11) + fig.tight_layout() + fig.savefig(OUT / "fig_ladder.png", dpi=140, facecolor="white", + bbox_inches="tight") + print("wrote fig_ladder.png") + + +# ---------------- fig 3: cross-scale grid ---------------- +def fig_scale(): + N2 = ROOT / "results-node2-final/results-12b" + N1 = ROOT / "results-node-final/results-12b" + panels = { + ("MBPP", "E2B"): [ + ("base", *ev(OUT / "eval_code_code_s0_full.json", 0)), + ("loop k=4", *ev(OUT / "eval_code_code_s0_full.json", 4)), + ("adaptive k=4", *ev(OUT / "eval_code_e2b_adaptive.json", 4)), + ("distill", *ev(OUT / "eval_code_distill_s1.json", 1)), + ], + ("MBPP", "12B"): [ + ("base", *ev(N1 / "eval_code_12b_trained.json", 0)), + ("loop k=4 (α=.3)", *ev(N1 / "eval_code_12b_trained.json", 4)), + ("adaptive k=4", *ev(N2 / "eval_code_12b_adaptive.json", 4)), + ("distill", *ev(OUT / "eval_code_12b_distill.json", 1)), + ], + ("GSM8K", "E2B"): [ + ("base", *ev(OUT / "eval_uni.json", 0)), + ("loop k=2", *ev(OUT / "eval_uni.json", 2)), + ("adaptive k=2", *ev(OUT / "eval_gsm_e2b_adaptive.json", 2)), + ("distill", *ev(OUT / "eval_gsm_distill_retry.json", 1)), + ], + ("GSM8K", "12B"): [ + ("base", *ev(N1 / "eval_12b_gsm_loop.json", 0)), + ("loop k=2 (α=.3)", *ev(N1 / "eval_12b_gsm_loop.json", 2)), + ("adaptive k=2", *ev(OUT / "eval_12b_gsm_adaptive.json", 2)), + ("distill*", *ev(OUT / "eval_12b_gsm_distill.json", 1)), + ], + } + fig, axes = plt.subplots(2, 2, figsize=(9.6, 6.6)) + for ax, ((task, scale), arms) in zip(axes.flat, panels.items()): + if (task, scale) == ("GSM8K", "12B"): + ax.annotate("*training collapse (0.00)", (3, 0.03), fontsize=7.5, + ha="center", color=RED) + style(ax) + x = range(len(arms)) + ax.bar([i - 0.19 for i in x], [a[1] for a in arms], width=0.36, + color=GRAY, alpha=0.85, label="overall") + ax.bar([i + 0.19 for i in x], [a[2] for a in arms], width=0.36, + color=BLUE, alpha=0.85, label="hard") + ax.axhline(arms[0][1], color=GRAY, lw=1, ls=":") + ax.set_xticks(list(x)) + ax.set_xticklabels([a[0] for a in arms], fontsize=8) + ax.set_title(f"{task} · {scale}", fontsize=10) + ax.set_ylim(0, 1.0) + if ax is axes.flat[0]: + ax.legend(fontsize=8, frameon=False) + fig.suptitle("Cross-scale attribution: constant-α destroys the 12B " + "substrate; state-dependent α restores it (MBPP) but not " + "everywhere", fontsize=11.5) + fig.tight_layout(rect=(0, 0, 1, 0.96)) + fig.savefig(OUT / "fig_scale.png", dpi=140, facecolor="white", + bbox_inches="tight") + print("wrote fig_scale.png") + + +# ---------------- fig 4: transfer panel ---------------- +def fig_transfer(): + def he_hard(path, k): + d = json.load(open(OUT / path)) + return d["ks"][str(k)]["acc"], d["ks"][str(k)]["hard_acc"] + + groups = [ + ("HumanEval\n(MBPP-trained)", [ + ("base", *he_hard("eval_humaneval_trained.json", 0)), + ("loop k=4", *he_hard("eval_humaneval_trained.json", 4)), + ("distill", *he_hard("eval_humaneval_distill_transfer.json", 1)), + ]), + ("Rust / MultiPL-E\n(Python-trained)", [ + ("base", *ev(OUT / "eval_rust_py_transfer.json", 0)), + ("loop k=4", *ev(OUT / "eval_rust_py_transfer.json", 4)), + ]), + ] + fig, axes = plt.subplots(1, 2, figsize=(7.6, 3.8)) + for ax, (title, arms) in zip(axes, groups): + style(ax) + x = range(len(arms)) + ax.bar([i - 0.19 for i in x], [a[1] for a in arms], width=0.36, + color=GRAY, alpha=0.85, label="overall") + ax.bar([i + 0.19 for i in x], + [a[2] if a[2] is not None else 0 for a in arms], + width=0.36, color=BLUE, alpha=0.85, label="hard") + ax.set_xticks(list(x)) + ax.set_xticklabels([a[0] for a in arms], fontsize=8.5) + ax.set_title(title, fontsize=9.5) + ax.set_ylim(0, 1.0) + axes[0].legend(fontsize=8, frameon=False) + fig.suptitle("Transfer: the implant moves with the substrate, " + "not the task", fontsize=11.5) + fig.tight_layout(rect=(0, 0, 1, 0.94)) + fig.savefig(OUT / "fig_transfer.png", dpi=140, facecolor="white", + bbox_inches="tight") + print("wrote fig_transfer.png") + + +if __name__ == "__main__": + fig_placement() + fig_ladder() + fig_scale() + fig_transfer() diff --git a/scripts/stats_final.py b/scripts/stats_final.py new file mode 100644 index 0000000..c6ecd64 --- /dev/null +++ b/scripts/stats_final.py @@ -0,0 +1,261 @@ +"""Final statistics pass over all eval per_item logs (CPU only). + +Produces results-loop/STATS.md + stats_final.json: + 1. Wilson 95% CIs for every headline number (overall + hard bucket). + 2. Exact McNemar tests for the key paired comparisons (same items). + 3. Pooled hard bucket across MBPP + HumanEval + Rust (paired k>0 vs k0 + within each item, pooled counts). + 4. Label-robustness check: hard bucket redefined via consensus k0 + failure across seeds instead of the single greedy labeling run. +""" + +import json +import math +from pathlib import Path + +OUT = Path(__file__).resolve().parent.parent / "results-loop" + + +def wilson(c, n, z=1.96): + if n == 0: + return (float("nan"), float("nan")) + p = c / n + d = 1 + z * z / n + ctr = (p + z * z / (2 * n)) / d + hw = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d + return (ctr - hw, ctr + hw) + + +def binom_two_sided(k, n): + """Exact two-sided binomial test p-value, p0=0.5 (for McNemar).""" + if n == 0: + return 1.0 + def pmf(i): + return math.comb(n, i) * 0.5 ** n + pk = pmf(k) + return min(1.0, sum(pmf(i) for i in range(n + 1) if pmf(i) <= pk + 1e-12)) + + +def mcnemar(pairs): + """pairs: list of (a_ok, b_ok). Returns dict with discordants + p.""" + b01 = sum(1 for a, b in pairs if not a and b) # b wins + b10 = sum(1 for a, b in pairs if a and not b) # a wins + return {"n": len(pairs), "a_only": b10, "b_only": b01, + "p": binom_two_sided(min(b01, b10), b01 + b10)} + + +def load_labels(path, key, lab_key="label"): + data = json.load(open(OUT / path)) + return {it[key]: it[lab_key] for it in data + if it.get("split", "test") == "test"} + + +def per_item(fname, k): + d = json.load(open(OUT / fname)) + v = d["ks"][str(k)] + key = "task_id" if "task_id" in v["per_item"][0] else "idx" + return {it[key]: it["ok"] for it in v["per_item"]} + + +def acc_ci(ok_map, subset=None): + ids = [i for i in ok_map if subset is None or i in subset] + c = sum(ok_map[i] for i in ids) + lo, hi = wilson(c, len(ids)) + return {"acc": c / len(ids) if ids else float("nan"), "n": len(ids), + "ci": [round(lo, 3), round(hi, 3)]} + + +def fmt(r): + return f"{r['acc']:.3f} [{r['ci'][0]:.3f}, {r['ci'][1]:.3f}] (n={r['n']})" + + +def main(): + mbpp_lab = load_labels("mbpp_data.json", "task_id") + he_lab = {k: ("hard" if v else "easy") # plan-reachable flag; hard needs k0-fail + for k, v in json.load(open(OUT / "humaneval_labels.json")).items()} + rust_lab = load_labels("rust_data.json", "task_id") + + mbpp_hard = {t for t, l in mbpp_lab.items() if l == "hard"} + rust_hard = {t for t, l in rust_lab.items() if l == "hard"} + + report = {} + lines = ["# Final statistics pass", ""] + + # ---------- 1. headline numbers with Wilson CIs ---------- + lines += ["## Headline numbers (Wilson 95% CIs)", ""] + ARMS = [ + # (label, file, k, hard-subset) + ("MBPP loop s0 k=0 (base)", "eval_code_code_s0_full.json", 0, mbpp_hard), + ("MBPP loop s0 k=2", "eval_code_code_s0_full.json", 2, mbpp_hard), + ("MBPP loop s0 k=4", "eval_code_code_s0_full.json", 4, mbpp_hard), + ("MBPP distill s1 k=1 (FF)", "eval_code_distill_s1.json", 1, mbpp_hard), + ("MBPP pause16 k=1", "eval_code_pause16.json", 1, mbpp_hard), + ("MBPP stack-train k=4", "eval_code_stack_train.json", 4, mbpp_hard), + ("MBPP distill-in-loopmode k=2", "eval_code_distill_loopmode.json", 2, mbpp_hard), + ("Rust transfer k=0", "eval_rust_py_transfer.json", 0, rust_hard), + ("Rust transfer k=4", "eval_rust_py_transfer.json", 4, rust_hard), + ] + arm_maps = {} + for label, f, k, hard in ARMS: + m = per_item(f, k) + arm_maps[label] = (m, hard) + o, h = acc_ci(m), acc_ci(m, hard) + report[label] = {"overall": o, "hard": h} + lines.append(f"- **{label}**: overall {fmt(o)}; hard {fmt(h)}") + + # HumanEval: hard = plan-reachable AND k0-fail (per its eval definition) + he_tr0 = per_item("eval_humaneval_trained.json", 0) + he_hard = {t for t in he_tr0 if he_lab.get(t) == "hard" and not he_tr0[t]} + for label, f, k in [("HumanEval loop k=4", "eval_humaneval_trained.json", 4), + ("HumanEval distill k=1", "eval_humaneval_distill_transfer.json", 1)]: + m = per_item(f, k) + o, h = acc_ci(m), acc_ci(m, he_hard) + arm_maps[label] = (m, he_hard) + report[label] = {"overall": o, "hard": h} + lines.append(f"- **{label}**: overall {fmt(o)}; hard {fmt(h)}") + + # best-of-3 / budget-CoT (no per_item; CI from counts) + for label, f in [("MBPP best-of-3 (compute-matched)", "eval_bestof3.json"), + ("MBPP budget-CoT-50", "eval_budgetcot.json")]: + d = json.load(open(OUT / f)) + n, nh = 500, len(mbpp_hard) + o = {"acc": d["acc"], "n": n, + "ci": [round(x, 3) for x in wilson(round(d["acc"] * n), n)]} + hacc = d["by_label"]["hard"] + h = {"acc": hacc, "n": nh, + "ci": [round(x, 3) for x in wilson(round(hacc * nh), nh)]} + report[label] = {"overall": o, "hard": h} + lines.append(f"- **{label}**: overall {fmt(o)}; hard {fmt(h)}") + + # distill seed spread + hs, os_ = [], [] + files = ["eval_code_code_distill.json"] + [ + f"eval_code_distill_s{s}.json" for s in range(1, 8)] + for f in files: + try: + m = per_item(f, 1) + except FileNotFoundError: + continue + hs.append(acc_ci(m, mbpp_hard)["acc"]) + os_.append(acc_ci(m)["acc"]) + mean = sum(hs) / len(hs) + sd = (sum((x - mean) ** 2 for x in hs) / (len(hs) - 1)) ** 0.5 + lines.append(f"- **MBPP distill, {len(hs)} runs (hard)**: mean {mean:.3f} " + f"± {sd:.3f} sd (range {min(hs):.3f}-{max(hs):.3f}); " + f"overall mean {sum(os_)/len(os_):.3f}") + report["distill_seed_spread"] = {"hard": hs, "overall": os_} + + # loop seed spread (s0..s4 k=4) + lh = [] + for tag in ["s0_full", "s1", "s2", "s3", "s4"]: + try: + m = per_item(f"eval_code_code_{tag}.json", 4) + lh.append(acc_ci(m, mbpp_hard)["acc"]) + except (FileNotFoundError, KeyError): + pass + if lh: + mean = sum(lh) / len(lh) + sd = (sum((x - mean) ** 2 for x in lh) / max(1, len(lh) - 1)) ** 0.5 + lines.append(f"- **MBPP loop seeds k=4 (hard)**: mean {mean:.3f} " + f"± {sd:.3f} sd (n_seeds={len(lh)})") + report["loop_seed_spread_hard"] = lh + + # ---------- 2. McNemar paired tests ---------- + lines += ["", "## McNemar exact tests (paired on items)", ""] + + def pair(m_a, m_b, subset=None): + ids = [i for i in m_a if i in m_b + and (subset is None or i in subset)] + return [(m_a[i], m_b[i]) for i in ids] + + loop0, _ = arm_maps["MBPP loop s0 k=0 (base)"] + loop4, _ = arm_maps["MBPP loop s0 k=4"] + dist1, _ = arm_maps["MBPP distill s1 k=1 (FF)"] + stack4, _ = arm_maps["MBPP stack-train k=4"] + + TESTS = [ + ("loop k=4 vs k=0, overall", loop0, loop4, None), + ("loop k=4 vs k=0, hard", loop0, loop4, mbpp_hard), + ("distill k=1 vs loop k=4, overall", loop4, dist1, None), + ("distill k=1 vs loop k=4, hard", loop4, dist1, mbpp_hard), + ("stack-train k=4 vs distill k=1, hard", dist1, stack4, mbpp_hard), + ("HumanEval loop k=4 vs k=0, overall", he_tr0, + arm_maps["HumanEval loop k=4"][0], None), + ("HumanEval distill k=1 vs k=0, overall", he_tr0, + arm_maps["HumanEval distill k=1"][0], None), + ] + rust0 = arm_maps["Rust transfer k=0"][0] + rust4 = arm_maps["Rust transfer k=4"][0] + TESTS += [("Rust loop k=4 vs k=0, overall", rust0, rust4, None), + ("Rust loop k=4 vs k=0, hard", rust0, rust4, rust_hard)] + + report["mcnemar"] = {} + for name, a, b, subset in TESTS: + r = mcnemar(pair(a, b, subset)) + report["mcnemar"][name] = r + sig = "**significant**" if r["p"] < 0.05 else "n.s." + lines.append(f"- {name}: A-only {r['a_only']}, B-only {r['b_only']}, " + f"n={r['n']}, p={r['p']:.4g} ({sig})") + + # ---------- 3. pooled hard bucket across benchmarks ---------- + lines += ["", "## Pooled hard bucket (MBPP + HumanEval + Rust)", + "", "Paired within-item k>0 vs k=0, counts pooled across " + "benchmarks (loop arm; distill pooled where available).", ""] + pooled_loop = (pair(loop0, loop4, mbpp_hard) + + pair(he_tr0, arm_maps["HumanEval loop k=4"][0], he_hard) + + pair(rust0, rust4, rust_hard)) + r = mcnemar(pooled_loop) + c_base = sum(a for a, _ in pooled_loop) + c_loop = sum(b for _, b in pooled_loop) + n = len(pooled_loop) + lines.append(f"- **loop**: base {c_base}/{n} -> loop {c_loop}/{n} " + f"({c_base/n:.3f} -> {c_loop/n:.3f}, CI " + f"{[round(x,3) for x in wilson(c_loop, n)]}), " + f"McNemar p={r['p']:.3g}") + report["pooled_hard_loop"] = {"base": c_base, "arm": c_loop, "n": n, + "mcnemar": r} + pooled_dist = (pair(loop0, dist1, mbpp_hard) + + pair(he_tr0, arm_maps["HumanEval distill k=1"][0], he_hard)) + r = mcnemar(pooled_dist) + c_base = sum(a for a, _ in pooled_dist) + c_d = sum(b for _, b in pooled_dist) + n = len(pooled_dist) + lines.append(f"- **distill**: base {c_base}/{n} -> distill {c_d}/{n} " + f"({c_base/n:.3f} -> {c_d/n:.3f}, CI " + f"{[round(x,3) for x in wilson(c_d, n)]}), " + f"McNemar p={r['p']:.3g}") + report["pooled_hard_distill"] = {"base": c_base, "arm": c_d, "n": n, + "mcnemar": r} + + # ---------- 4. label robustness: consensus-k0 hard set ---------- + lines += ["", "## Label robustness (consensus-k0 hard set)", "", + "Hard bucket redefined as: labeled hard AND k=0 fails in " + "every seed's own eval run (removes single-greedy-run " + "selection noise).", ""] + k0maps = [] + for tag in ["s0_full", "s1", "s2", "s3", "s4"]: + try: + k0maps.append(per_item(f"eval_code_code_{tag}.json", 0)) + except (FileNotFoundError, KeyError): + pass + consensus = {t for t in mbpp_hard + if all(not m.get(t, False) for m in k0maps)} + lines.append(f"- consensus hard set: {len(consensus)} of " + f"{len(mbpp_hard)} labeled-hard items") + for label in ["MBPP loop s0 k=4", "MBPP distill s1 k=1 (FF)", + "MBPP stack-train k=4"]: + m, _ = arm_maps[label] + r0 = acc_ci(m, mbpp_hard) + rc = acc_ci(m, consensus) + lines.append(f"- {label}: labeled-hard {r0['acc']:.3f} -> " + f"consensus-hard {fmt(rc)}") + report.setdefault("consensus_hard", {})[label] = rc + + (OUT / "STATS.md").write_text("\n".join(lines) + "\n") + json.dump(report, open(OUT / "stats_final.json", "w"), indent=1) + print("\n".join(lines)) + print("\nwrote", OUT / "STATS.md", "and stats_final.json") + + +if __name__ == "__main__": + main()