PLAN_SELFPACED.md: v2 prototype plan for learned workspace-compute gating (E0 numbers grounded, E1-E3 pre-registered)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Nils
2026-07-15 22:50:35 +02:00
co-authored by Claude Fable 5
parent 600f965247
commit 5f58bc8aed
2 changed files with 82 additions and 0 deletions
+1
View File
@@ -13,6 +13,7 @@ band; what does it actually buy?). The second is the active one.
| What worked / what failed / design rules / ops pitfalls | `LESSONS.md` — read before running anything on this hardware | | What worked / what failed / design rules / ops pitfalls | `LESSONS.md` — read before running anything on this hardware |
| Pre-registrations + scored outcomes (17 items, incl. refutations) | `results-loop/PROTOCOL_UNIFIED.md` — the methods backbone; every claim in PAPER.md §3 traces to an item here | | Pre-registrations + scored outcomes (17 items, incl. refutations) | `results-loop/PROTOCOL_UNIFIED.md` — the methods backbone; every claim in PAPER.md §3 traces to an item here |
| Significance tests behind any claimed number | `results-loop/STATS.md` | | Significance tests behind any claimed number | `results-loop/STATS.md` |
| The v2 prototype plan (self-paced workspace: learned gating) | `PLAN_SELFPACED.md` |
| Lab-notebook narrative of the looping investigation | `WORKSPACE_LOOPING.md` (superseded where it disagrees with PAPER.md) | | Lab-notebook narrative of the looping investigation | `WORKSPACE_LOOPING.md` (superseded where it disagrees with PAPER.md) |
| Base-reproduction results (lens replication itself) | `RESULTS.md`, `README.md` | | Base-reproduction results (lens replication itself) | `RESULTS.md`, `README.md` |
| Per-model lens maps: workspace bands, KV-share boundaries, pinned revisions | `results/REGIMES.json` (canonical registry) + `results/jbar*.pt` (raw J̄) + `results/exp4*.log` (regime scans) | | Per-model lens maps: workspace bands, KV-share boundaries, pinned revisions | `results/REGIMES.json` (canonical registry) + `results/jbar*.pt` (raw J̄) + `results/exp4*.log` (regime scans) |
+81
View File
@@ -0,0 +1,81 @@
# Prototype plan: the self-paced workspace (v2)
*Drafted 2026-07-16, pre-registration-style. Goal: test whether the model
can learn to allocate workspace-loop compute ON ITS OWN — per prompt and
per generation step — rather than at a swept hyperparameter k.*
## The concept
At every step the system chooses: emit, or spend a band iteration updating
the workspace first. Make that choice a learned gate g(workspace state).
Compute becomes a decision, not a constant. Gate-bought iterations emit no
tokens, so they are exposure-safe by construction (deterministic given
state — the pause-position property).
## What already exists (de-risked ingredients)
| ingredient | evidence | where |
|---|---|---|
| static per-prompt gate (E0) | gated 52.0 overall, easy 97.5 (vs 88.5 uniform), hard 28.6; bottleneck = probe recall (18/28 tp, 95 predicted hard) | `gate_probe.py`, `eval_gated.json` |
| state-dependent control heads train | adaptive-α rescued 12B (3.8K params) | AdaptiveMergeAdapter |
| evolving state during generation | carry beats registers on GSM (hard 0→9.4) | `carry_common.py`, `eval_carry.json` |
| loop-capacity knob | rung-2 band-LoRA = best hard numbers (42.9/46.4) | `lora_band.py` |
| gates over DEPTH (complementary axis) | learnable compute envelope g[t,l] | `path_gates.py` (Nils, in progress) |
| free gate supervision | STaR difficulty labels; per-position labels derivable | prep_star/prep_mbpp |
## Experiments
### E1 — learned per-prompt halting (prompt side, MBPP)
Replace fixed k with a trained soft halting gate. Architecture: after each
iteration i, gate head h(e, ŝ_i) → p_halt,i (zero-init to fixed-k
behavior); training uses the soft mixture of iteration outputs weighted by
halting distribution (ACT-style), CE + λ·E[iterations] compute penalty;
deploy = argmax halt. Trains end-to-end, NO RL. Arms: λ ∈ {1e-3, 1e-2},
vs E0 probe-gate and uniform-k anchors.
**Pre-registered predictions:** (a) accuracy ≥ uniform k=4 overall at ≤60%
of its mean iterations; (b) easy ≥ 95% (gate protects the substrate);
(c) allocation correlates with STaR label (point-biserial r > 0.3);
(d) hard ≥ E0's 28.6% (learned gate beats frozen probe recall).
**Failure mode to watch:** gate collapse (all-0/all-1) — mitigate with
penalty warmup + entropy bonus; collapse at all λ falsifies E1.
### E2 — generation-side gating (GSM, gated carry)
Substrate: design-C carry + short verified-CoT supervision (dense targets;
harvest with "solve in ≤3 short steps", answer-verified). Gate per token
step decides whether the carry state updates through the band or passes
through: x_t = g·merge(e_t, s_{t1}) + (1g)·e_t, penalty λ·E[g].
Anchors: carry-always, carry-never (same supervision).
**Predictions:** (a) gate fires non-uniformly, concentrated near numeric/
operator tokens (measurable); (b) accuracy ≥ carry-always (gating as
protection); (c) easy-bucket damage < carry-always's (83→45 was the
unprotected number). Hard-bucket *gain* over carry-always is hoped for,
not predicted.
### E3 — power knob (only if E1 or E2 shows clean gating)
Warm-start rung-2 band-LoRA under the gate; joint fine-tune. Question: do
gate-bought iterations do MORE per iteration with a trainable band?
Metric: the internalization count (how many scratchpad steps can be
removed post-hoc, E2 curriculum) as a function of LoRA rank.
### Lens verification (throughout — our home advantage)
J-lens reads of gated vs ungated positions: do bought iterations sharpen
task-relevant concepts at the positions where the gate fired? This is the
mechanistic check that the gate allocates *meaningfully*, not just
correlationally.
## Explicitly out of scope for the prototype
Outcome-RL training of the gate (GRPO with compute price) — stage 2, only
if E1E3 show selective gating. 12B/scale transfer. Cross-task gates.
## Budget & order
E1: 3 arms × ~75 min (Spark). E2: harvest ~30 min + 3 arms × ~90 min.
E3: +2 arms. Total ≈ 1.5 Spark-days. Runs after the lens campaign; queue
via gpuq as usual, every arm pre-registered in PROTOCOL_UNIFIED.md before
launch (items 18+).
## Kill criteria (decided in advance)
- E1 gate collapse at all λ AND E2 uniform firing → the state does not
carry usable "needs compute" signal at this scale; program stops, E0's
static-gate deployment note stands as the practical answer.
- E1 works but hard < E0 → learned gate worse than probe; ship probe-gate,
keep E2 only if its (a)/(b) hold.