# Results: J-lens reproduction on gemma-4-E2B-it (DGX Spark) Reproduction of the core claims of *"Verbalizable Representations Form a Global Workspace in Language Models"* (Transformer Circuits, 2026) on `google/gemma-4-E2B-it` (35 layers, d=1536, tied embeddings). The paper's subjects are Claude Sonnet/Haiku/Opus (≈50–130 layers); this tests transfer to a small open-weights model on local hardware. ![The residual stream under the J-lens: four depth regimes, and whether the workspace band can be looped](results/jlens_diagram.png) **Jacobian:** J̄_ℓ = E[∂h_final,t′/∂h_ℓ,t] averaged *exactly* over all (t, t′≥t) pairs of 300 fineweb-edu prompts (seq len 64) ≈ 620k Jacobian samples per layer; 1536 batched VJPs per prompt, ~21 s/prompt, ~105 min total. Correctness anchor: the last-layer pair-sum equals T·I exactly. Results were stable between the 100-prompt checkpoint and the full 300-prompt average. ## 1. Verbal report / internal reasoning readouts — REPRODUCES **Two-hop ("The animal that spins webs has how many legs?" → "8"):** The unspoken intermediate **spider** is read by the J-lens at (layer 8, " webs") P≈0.25, peaking at (layer 33, " has") P=0.39, and again at the answer position (layers 28–31), while candidate numbers (`two three four seven five six` → `Eight six`) appear at layers 24–28 just before output. See `results/heat_spider.png`. **Multilingual intermediates (小 → 大):** during the Chinese antonym task the lens reads English/multilingual abstractions mid-network — `opposite`, `contraire`, `反対`, then `bigger/larger/large/lớn/大き` at layers 24–28 — before the Chinese output token. Matches the paper's finding that intermediate computation is partly in a cross-lingual/English-leaning code. **Layer structure (paper: three regimes + workspace only in intermediate layers):** early layers (0–6) read as noise (no stable content); layers ~8–12 read task-frame concepts (`number`, `answer`, `insect`, `word`); layers ~24–31 read task *content* (spider, numbers, opposite/big); the final layers become a motor regime (top lens token = next output token 68% at L34 vs ~0% early). `results/layer_profile.png`. ## 2. Directed modulation ("hold citrus in mind") — WEAK/PARTIAL Instructed to think about citrus fruits while copying unrelated text, the model copies perfectly and the lens shows citrus concepts elevated ~30× over the control condition during generation (max P 0.011 vs 0.0003) — but at tiny absolute probability. The paper's strong "orange/lemon in the workspace throughout" does not appear at this scale; only a trace of it. ## 3. Causal swaps — REPRODUCES WITH ONE ADAPTATION The paper's swap `h ← h + V(σ(c)−c)`, `c = V⁺h` along J-lens vectors either had no effect (gated, few slots) or garbled generation (ungated, all positions) on this model. Diagnosis: J-lens vectors here are good **readers** but poor **writers** — the minimal-norm edit that swaps v_src·h ↔ v_tgt·h erases the source from the lens (P(France): 0.21 → 3e-18) without writing the target, because all J-lens vectors share dominant J̄ components (raw cosine France/China 0.74–0.92). Working variant (`jlens/core.py: swap_hooks`): **gate with the J-lens** (intervene only at (layer, position) slots where the lens reads the source concept, P>thr), but **write along tied token-embedding directions** (move h's projection on ê_src onto ê_tgt). With thr=0.01, α=1: - **Broadcast France→China** (paper §flexible generalization, their success ~40–53%): capital `Paris→Beijing` ✓, language `French→Chinese` ✓, river `Seine→Nile` ✓; continent/currency/food unchanged → **3/6 templates**. - **Capital grid over 6 countries (30 ordered pairs): 6/30 flips**, all among France/China/Japan(/Canada→France) — exactly the countries the lens reads strongly; for Egypt/Brazil/Canada as source the gate rarely fires (concept weakly loaded in the workspace). This mirrors the paper's finding that swap success tracks the concept's workspace loading, and their model-size dependence (54–70% on their much larger models). - **Two-hop spider→ant:** transfers the concept but breaks the format — output becomes `ant` (α=2) instead of the hoped-for `6`; reverse direction no effect (ant is weakly loaded: 2 gated slots vs spider's ~9). - Controls: no-swap and spider→piano leave the answer at `8`; the paper-basis write (`write="jlens"`) leaves it at `8`. ## 4. J-space capacity (paper: ~10% of variance, k≤25 active) — SHAPE ONLY Non-negative matching pursuit (k≤25 J-lens vectors, NNLS refit) explains **<1% of residual variance in intermediate layers** (peak 0.9% at L25) and **22% at the final layer**. Qualitatively consistent with a small privileged subspace distinct from the bulk of activation variance, but the mid-network occupancy is an order of magnitude below the paper's ~10% — either a real scale difference or a limitation of our 300-prompt J̄ estimate. ## 5. Sensor / workspace / motor regimes and ignition — REPRODUCES Per-layer diagnostics over 8 prompts (`scripts/exp4_regimes.py`, `results/regimes.png`): - **Sensor: layers ~6–13.** Top J-lens token equals the *current input* token at up to 26% of positions (exactly 0% everywhere after L13). Low persistence across positions — content changes with each token. - **Workspace: layers ~14–30.** Input echo vanishes abruptly at L14; cross-position persistence of lens content *peaks* (top-10 Jaccard ~0.2–0.26 at L14–21 vs ~0.05 in the sensor band) — content held stable across positions; this is also where abstract task content lives (spider, numbers, opposite/big at L19–31 in exp1) and where gated swaps fire. - **Motor: layers ~31–34.** Top lens token equals the *next output* token at 42–48% of positions; J-space R² jumps to 22%. **Parameters per regime** (decoder layers, of 1.86B total): early 0–5 = 224M (12%), sensor 6–13 = 297M (16%), workspace 14–30 = 1.08B (58%), motor 31–34 = 261M (14%). Notably the architecture itself pivots exactly at the observed sensor→workspace boundary: from layer 15 the model switches to double-wide MLPs with cross-layer-shared KV (63.7M/layer vs 36.2M/layer before), i.e. the designers put the parameter mass where the workspace content lives. (Non-decoder params: 403M tied embeddings, 2.36B per-layer embedding machinery, 476M vision/audio towers; 5.1B raw total.) **Ignition** (paper: ambiguous inputs produce sharp binary commitment at workspace onset, not proportional mixing): replacing one token's embedding (main + per-layer) with w·concept1 + (1−w)·concept2 and reading the lens at that position (`results/ignition.png`, `ignition_atpos.pt`): - Sensor layer 13 mixes **gradually**: e.g. ocean/violin commitment C = −0.92, −0.70, −0.41, −0.12, +0.36, +0.92, +1.00 across w = 0→1. - Deep layer 31 commits **all-or-nothing** at w≈0.5: −1.00, −1.00, −1.00, +0.97, +1.00 (dog/piano); same step shape for all three concept pairs. ## Verdict | Paper claim | This model | |---|---| | Unspoken intermediates readable via J-lens | ✓ strong (spider, big/opposite) | | Three depth regimes; workspace = intermediate layers | ✓ | | Cross-lingual abstract code mid-network | ✓ | | Directed modulation loads workspace | ~ trace only | | J-lens coordinate swaps flip downstream behavior | ✓ after adaptation (embedding write basis, J-lens gating); success tracks workspace loading | | Broadcast/flexible generalization of swapped concept | ✓ 3/6 templates | | J-space is small fraction of activation variance | ✓ shape (but <1% vs their ~10%) | | Ignition: binary commitment from workspace onset | ✓ (graded at L13, step function at L31) | Overall: the *readable global-workspace* phenomenology transfers remarkably well to a 2B-class open model; the *causal* properties transfer partially and require respecting the model's write basis (tied embeddings) rather than the J-lens basis itself. ## 6. Scale comparison: gemma-4-12B-it (48 layers, d=3840) Same pipeline, J̄ from 150 fineweb-edu prompts (177 s/prompt, ~7.4 h on the GB10). Logs/plots in `results-12b/`. **Readouts get much stronger with scale** (paper predicts this): | readout | E2B | 12B | |---|---|---| | unspoken "spider" peak P | 0.39 | **0.80** | | English "big/large" in Chinese task | 0.031 | **0.33** | | directed modulation (citrus vs control) | ~30× trace | none detected | **The write basis migrates with scale — the central new finding.** The gated swap that worked on E2B with *token-embedding* writes fails on 12B with the same erase-without-write signature the J-lens basis showed on E2B (France deleted — "Please specify which country" — but China never written). What works on 12B: **activation-derived concept vectors** (mean residual at the concept token over 4 contrastive templates, per layer) as the write direction, still gated by the J-lens reading. With that: - **Capital grid: 30/30** ordered country pairs flip correctly (E2B embedding-write: 6/30), α robust over 0.5–2.0. - **Broadcast 4/4**: Paris→Beijing, French→Chinese, Europe→Asia, and Seine→**Yangtze** (the correct Chinese river, where E2B gave generic "Nile"). Exceeds the paper's reported broadcast rates (~40–53%). - Two-hop spider→ant still does not flip the "8" at 12B (attribute lookup appears to route around the gated slots). Summary of write bases: J-lens vectors read everywhere but never write on gemma (unlike the paper's Claude models); token embeddings write at 2B (small tied model keeps concepts embedding-aligned); by 12B concepts have rotated into model-specific activation directions. The J-lens *gate* (where/when to intervene) transfers across scales unchanged. **Regime geometry shifts deeper (fractionally):** sensor echo spans L11–35 (peak 0.36 at L25, ≈0.23–0.73 fractional vs 0.17–0.40 on E2B); abstract workspace content (spider P=0.80, swap gates) concentrates at L36–45 (≈0.75–0.94); motor alignment ramps only in the last few layers and is weaker under our teacher-forced metric (0.12 vs 0.48). `results-12b/regimes.png`. **J-space occupancy** rounds to 0.000 at all sampled intermediate layers (k≤25 pursuit; even sparser than E2B's <1%). The final-layer slot was not sampled at 12B (stride artifact). **Ignition**: the clean graded-vs-binary contrast we measured on E2B did not replicate cleanly on 12B with the at-position protocol — mixed-token readings there are dominated by contextual priors rather than the injected embedding. Would need the paper's exact protocol to adjudicate. ## 7. MoE: gemma-4-26B-A4B-it (30 layers, d=2816, 128 experts, top-8, ~4B active) Same pipeline, J̄ from 100 fineweb-edu prompts. The exact-Jacobian method survives the router untouched: the vmapped backward traverses the top-8 expert gather/scatter with **0.0 last-layer identity error** — averaging Jacobians over prompts that each route through *different* experts still yields a coherent lens. Cost 487 s/prompt (~13.5 h), ~4× the active-param prediction because the backward decomposes into scattered per-expert matmuls. Logs/plots in `results-26b/`. **Readouts are the strongest of all four models:** | readout | E2B | 12B | 26B-MoE | |---|---|---|---| | unspoken "spider" peak P | 0.39 | 0.80 | **1.00** | | English "big/large" in Chinese task | 0.031 | 0.33 | 0.36 | **Directed modulation finally works — and only here.** "Hold citrus in mind while copying unrelated text": citrus concepts reach **P=0.69** in the J-lens *during generation* (mean 0.005) vs 0.006 in the control — the paper's "orange/lemon held in the workspace throughout" result, which was a bare trace at E2B and undetectable at 12B. The routed model is the first to show it clearly, plausibly because a dedicated expert can maintain the held concept without disrupting the copy stream. **Write basis matches 12B:** embedding-write swap fails (0/30 grid, garbled output); **activation-derived concept vectors give 30/30** on the capital grid, J-lens-gated. Confirms the 12B finding that above ~10B, mid-network concepts leave the embedding basis — and that this is about scale, not the dense/MoE distinction. **Regimes are weaker and flatter** (`results-26b/regimes.png`): sensor echo peaks at just 0.13 (L10–19) vs 0.26 (E2B) / 0.36 (12B); workspace content L20–27; motor only the last 2 layers. Routing appears to spread the workspace across experts, so no single J-lens direction dominates any slot — consistent with **J-space R² ≈ 0.000 at every sampled layer** (even sparser than the dense models). The workspace is present (readouts prove it) but diffuse. ## Four-model summary | | E2B | 12B | 26B-A4B (MoE) | 31B | |---|---|---|---|---| | layers / d_model | 35 / 1536 | 48 / 3840 | 30 / 2816 | 60 / 5376 | | J̄ prompts | 300 | 150 | 100 | — | | Jacobian s/prompt | 21 | 177 | 487 | ~500 (est) | | unspoken "spider" P | 0.39 | 0.80 | 1.00 | — | | directed modulation | trace | none | **P=0.69** | — | | swap write basis | embedding | activation | activation | — | | capital grid | 6/30 | 30/30 | 30/30 | — | | J-lens gate (where) | ✓ | ✓ | ✓ | — | Consistent story across scale and architecture: **readouts strengthen with scale**; the **J-lens gate** (where to intervene) is universal; the **write basis migrates** from token-embedding (2B) to model-specific activation directions (≥12B, dense or MoE); the MoE keeps a workspace but a diffuse one. The one property that appeared only at the MoE — robust directed modulation — is the paper's clearest "deliberate control" signature. ## Repro notes - Files: see README. Full logs: `results/exp{1,2,3}.log`; plots `results/*.png`; averaged Jacobian `results/jbar.pt` (330 MB, fp32). - Total compute: ~2 h on the GB10 for J̄ + ~30 min for all experiments. - Caveats: single-token concepts only (paper limitation too); J̄ from 300 prompts at seq len 64 (paper: ~1000 prompts, unknown length); bf16 gradients (independent-run correlation of J̄ entries ≥0.996 layer 5+).