HumanEval transfer: trained-vs-untrained-merge paired test (p=0.80, n.s.) -- transfer gain is the merge itself, not trained content
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -278,10 +278,13 @@ changes with scale, and per-task tuning remains unavoidable.
|
||||

|
||||
|
||||
MBPP-trained implants applied unchanged: **HumanEval** overall 58.5%→66.5%
|
||||
(loop k=4, p=0.011; hard 0→31.6%); notably the *untrained* merge already
|
||||
reaches 64.6% and the transferred pause adapter 66.5% (hard 38.9%) — the
|
||||
transfer is substrate-shaped (a generically useful perturbation+content
|
||||
mode), not task-memorized. **Rust/MultiPL-E** (Python-trained, different
|
||||
(loop k=4, p=0.011 vs base; hard 0→31.6%). The decisive control: the
|
||||
*untrained* merge already reaches 64.6%, and trained-vs-untrained is **not
|
||||
significant** (paired McNemar at k=2, 9 vs 7 discordant, p=0.80). What
|
||||
transfers significantly is the *merge perturbation itself*, not the
|
||||
MBPP-trained content — the cleanest evidence that off-distribution value is
|
||||
substrate-shaped rather than task-memorized. (The transferred pause adapter
|
||||
reaches 66.5%, hard 38.9%, consistent with the same reading.) **Rust/MultiPL-E** (Python-trained, different
|
||||
language, compile-run-verified): hard 8.0%→24.0% (p=0.125 at n=25 —
|
||||
directionally consistent, underpowered). **Blocksworld** MBPP-transfer:
|
||||
hard 0→14.3% (task-trained: 43%). Content transfers where the substrate's
|
||||
|
||||
Reference in New Issue
Block a user