Files
jspace/results-loop/STATS.md
T

4.5 KiB

Final statistics pass

Headline numbers (Wilson 95% CIs)

  • MBPP loop s0 k=0 (base): overall 0.518 [0.474, 0.561] (n=500); hard 0.055 [0.019, 0.149] (n=55)
  • MBPP loop s0 k=2: overall 0.530 [0.486, 0.573] (n=500); hard 0.309 [0.203, 0.440] (n=55)
  • MBPP loop s0 k=4: overall 0.536 [0.492, 0.579] (n=500); hard 0.436 [0.314, 0.567] (n=55)
  • MBPP distill s1 k=1 (FF): overall 0.544 [0.500, 0.587] (n=500); hard 0.418 [0.297, 0.550] (n=55)
  • MBPP pause16 k=1: overall 0.552 [0.508, 0.595] (n=500); hard 0.364 [0.249, 0.496] (n=55)
  • MBPP stack-train k=4: overall 0.524 [0.480, 0.567] (n=500); hard 0.345 [0.234, 0.477] (n=55)
  • MBPP distill-in-loopmode k=2: overall 0.458 [0.415, 0.502] (n=500); hard 0.200 [0.116, 0.324] (n=55)
  • Rust transfer k=0: overall 0.591 [0.512, 0.665] (n=154); hard 0.080 [0.022, 0.250] (n=25)
  • Rust transfer k=4: overall 0.552 [0.473, 0.628] (n=154); hard 0.240 [0.115, 0.434] (n=25)
  • HumanEval loop k=4: overall 0.665 [0.589, 0.732] (n=164); hard 0.316 [0.191, 0.475] (n=38)
  • HumanEval distill k=1: overall 0.646 [0.571, 0.715] (n=164); hard 0.237 [0.130, 0.392] (n=38)
  • MBPP best-of-3 (compute-matched): overall 0.572 [0.528, 0.615] (n=500); hard 0.327 [0.218, 0.459] (n=55)
  • MBPP budget-CoT-50: overall 0.538 [0.494, 0.581] (n=500); hard 0.400 [0.281, 0.532] (n=55)
  • MBPP distill, 8 runs (hard): mean 0.457 ± 0.046 sd (range 0.400-0.545); overall mean 0.555
  • MBPP loop seeds k=4 (hard): mean 0.375 ± 0.055 sd (n_seeds=5)

McNemar exact tests (paired on items)

  • loop k=4 vs k=0, overall: A-only 30, B-only 39, n=500, p=0.3356 (n.s.)
  • loop k=4 vs k=0, hard: A-only 1, B-only 22, n=55, p=5.722e-06 (significant)
  • distill k=1 vs loop k=4, overall: A-only 33, B-only 37, n=500, p=0.7202 (n.s.)
  • distill k=1 vs loop k=4, hard: A-only 10, B-only 9, n=55, p=1 (n.s.)
  • stack-train k=4 vs distill k=1, hard: A-only 10, B-only 6, n=55, p=0.4545 (n.s.)
  • HumanEval loop k=4 vs k=0, overall: A-only 5, B-only 18, n=164, p=0.01062 (significant)
  • HumanEval distill k=1 vs k=0, overall: A-only 7, B-only 17, n=164, p=0.06391 (n.s.)
  • HumanEval trained vs UNTRAINED merge (k=2), overall: A-only 7, B-only 9, n=164, p=0.8036 (n.s.)
  • Rust loop k=4 vs k=0, overall: A-only 11, B-only 5, n=154, p=0.2101 (n.s.)
  • Rust loop k=4 vs k=0, hard: A-only 0, B-only 4, n=25, p=0.125 (n.s.)

Token baselines (paired, per-item)

  • best-of-3 ORACLE (any-pass): overall 0.578 [0.534, 0.621] (n=500); hard 0.345 [0.234, 0.477] (n=55)
  • best-of-3 oracle (selector run): overall 0.578 [0.534, 0.621] (n=500); hard 0.345 [0.234, 0.477] (n=55)
  • best-of-3 DEPLOYABLE (logprob-selected): overall 0.550 [0.506, 0.593] (n=500); hard 0.273 [0.173, 0.402] (n=55)
  • bo3-oracle vs loop k=4, overall: arm-only 27, bo3-only 48, p=0.0203 (significant)
  • bo3-oracle vs loop k=4, hard: arm-only 14, bo3-only 9, p=0.4049 (n.s.)
  • bo3-oracle vs distill k=1, overall: arm-only 27, bo3-only 44, p=0.05681 (n.s.)
  • bo3-oracle vs distill k=1, hard: arm-only 14, bo3-only 10, p=0.5413 (n.s.)
  • bo3-deployable vs loop k=4, overall: arm-only 31, bo3-only 38, p=0.4704 (n.s.)
  • bo3-deployable vs loop k=4, hard: arm-only 16, bo3-only 7, p=0.09314 (n.s.)
  • bo3-deployable vs distill k=1, overall: arm-only 31, bo3-only 34, p=0.8043 (n.s.)
  • bo3-deployable vs distill k=1, hard: arm-only 15, bo3-only 7, p=0.1338 (n.s.)
  • budget-CoT-50 (per-item rerun): overall 0.538 [0.494, 0.581] (n=500); hard 0.382 [0.265, 0.514] (n=55)
  • budget-CoT vs loop k=4, overall: arm-only 36, cot-only 37, p=1 (n.s.)
  • budget-CoT vs loop k=4, hard: arm-only 14, cot-only 11, p=0.69 (n.s.)
  • budget-CoT vs distill k=1, overall: arm-only 29, cot-only 26, p=0.7877 (n.s.)
  • budget-CoT vs distill k=1, hard: arm-only 8, cot-only 6, p=0.7905 (n.s.)

Pooled hard bucket (MBPP + HumanEval + Rust)

Paired within-item k>0 vs k=0, counts pooled across benchmarks (loop arm; distill pooled where available).

  • loop: base 5/118 -> loop 42/118 (0.042 -> 0.356, CI [0.275, 0.446]), McNemar p=1.46e-10
  • distill: base 3/93 -> distill 32/93 (0.032 -> 0.344, CI [0.255, 0.445]), McNemar p=2.98e-08

Label robustness (consensus-k0 hard set)

Hard bucket redefined as: labeled hard AND k=0 fails in every seed's own eval run (removes single-greedy-run selection noise).

  • consensus hard set: 52 of 55 labeled-hard items
  • MBPP loop s0 k=4: labeled-hard 0.436 -> consensus-hard 0.423 [0.299, 0.558] (n=52)
  • MBPP distill s1 k=1 (FF): labeled-hard 0.418 -> consensus-hard 0.404 [0.282, 0.539] (n=52)
  • MBPP stack-train k=4: labeled-hard 0.345 -> consensus-hard 0.346 [0.232, 0.482] (n=52)