budget-CoT per-item: paired ties with latent arms confirmed (p>=0.69)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -247,7 +247,10 @@ samples, so the comparison is exact), best-of-3 is statistically
|
||||
indistinguishable from the latent arms overall, *behind* them
|
||||
directionally on the hard bucket, and pays ≈3× visible tokens and serial
|
||||
decode for it. Budget-CoT-50 remains the strongest honest token baseline
|
||||
(53.8% overall, hard 40.0%). The implant's advantages at matched compute
|
||||
(53.8% overall, hard 40.0%; per-item rerun 53.8/38.2) — and the paired
|
||||
tests confirm it is a *tie* with the latent arms on both axes (p≥0.69 vs
|
||||
loop and distill), at the price of 50 visible tokens and their serial
|
||||
decode latency. The implant's advantages at matched compute
|
||||
without a verifier: zero visible tokens, zero decode overhead, and the
|
||||
hard-slice crown under distillation (46%). Where a task *does* come with a
|
||||
cheap verifier, oracle-style sampling is the better overall-accuracy
|
||||
|
||||
Reference in New Issue
Block a user