Repository navigation
fix(research): eleven verdicts were windows pretending to be checkpoints - #563
Merged
Merged
Conversation
campaignB_stats.row() did `np.concatenate([dvec(D, m, arm, ref) for m in
models])` and handed paired() 140 windows from four models as 140 replicates.
Windows replicate the TEXT, not the model family. This is the sixth instance of
this error in the campaign and the first found in a script rather than in prose
-- the previous five were caught by re-reading documents; this one was still
executing.
row() now takes n = models for cross-model claims, each contributing its own
mean log-ratio, and keeps windows only for single-model rows, which are
within-model claims entitled to them. Eleven of fourteen verdicts flip: every
"BEATS MXFP4" and every "loses to NF4" becomes a TIE. Point estimates barely
move (-4.99% -> -4.76%); the intervals grow by ~sqrt(35) and all now contain
zero. The correction is symmetric -- four rows moved toward our codebooks, seven
away -- which is how it is distinguishable from a motivated one.
Restated at the model level, in the documents that quoted the pooled figures:
- T40's decomposition: (+0.464%) x (-4.324%) = -3.880%, residual 6.94e-18. The
decomposition is arithmetic and holds at either level; what is withdrawn is
the strength of the surrounding margins. At four checkpoints "NF4 beats MXFP4"
is a TIE (-3.88%, p = 0.208) and NF4's win is a per-model result on 3 of 4.
- MX-asym-NEAR0's "-4.74% held-out, 4/4": the rotation is UNANIMOUS, so that
mean is algebraically identical to the plain four-model mean. The held-out
label transports nothing. Rotation-honest counterpart on the full nine-arm
pool: -2.46%, three distinct arms, so it is a protocol result with no codebook
to attach it to. "The only arm clearing significance" is false: four clear
uncorrected, none clears Bonferroni over the nine the argmin came from.
The assert sweep (187 sites) found the clipping-arm defect had an unfixed twin:
campaignB_books asserted max(abs(x)) under a docstring reading "BOTH tails", and
shipped TWO books at +1.000/-0.750 -- MX-asym-TOP and JK-asym-TOP, the second
never named. check() now verifies the kind label rather than trusting it: a
"clip" that does not clip fails too. The Bonferroni family is three placements,
not four; the reclassification was made on structural grounds a day before
anyone computed which way it moved a verdict, and at model level it moves none.
Corrects this repo's own CLIPPING_ARM_CORRECTION, whose blast-radius sentence
("campaignB_stats hard-codes its own list and never called check()") was true in
every word and used as an exemption when it was the opposite: not drawing from
candidates() is exactly why that campaign did not inherit the fix.
T41 is recorded as REFUTED with its literature check: the arm lowers squared
weight error by 6.8-11.3% on every checkpoint while raising perplexity on two,
and its "bit-exact" corollary was measured on torch.randn where the stated
exception cannot occur -- on real checkpoints it fails on 0.1-0.9% of blocks.
The seed-vs-distance control ran and did NOT break the confound; reported as
such, with the seed hypothesis's decisive evidence replicating 1 of 4.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
campaignB_stats.row()did this:Windows replicate the text, not the model family. Sixth instance of this error in the campaign, and the first found in a script rather than in prose — the previous five were caught by re-reading documents; this one was still executing. The section header said POOLED OVER ALL FOUR MODELS while the statistic said n = 140 windows.
Eleven of fourteen verdicts flip
…and seven more. The point estimates barely move (−4.99 % → −4.76 %); the intervals grow by ≈√35 and every one now contains zero. One verdict survives, and its own tag already says
3/4 in-sample.The correction is symmetric — four rows moved toward our codebooks, seven away. An error that only ever flattered us would be a different finding.
Consequences, restated in the documents that quoted the pooled figures
(+0.464 %) × (−4.324 %) = −3.880 %, residual6.94e-18. The composition is arithmetic and holds at either level; what is withdrawn is the strength of the surrounding margins. At four checkpoints "NF4 beats MXFP4" is a TIE (−3.88 %, p = 0.208).MX-asym-NEAR0's "−4.74 % held-out, 4/4": the rotation is unanimous, so that mean is algebraically identical to the plain four-model mean. The held-out label transports nothing. Rotation-honest counterpart on the full nine-arm pool: −2.46 %, three distinct arms — a protocol result with no codebook to attach it to. "The only arm clearing significance" is false: four clear uncorrected, none clears Bonferroni over the nine the argmin came from.The assert sweep (187 sites) found the clipping arm had a twin
campaignB_booksassertedmax(abs(x))under a docstring reading BOTH tails and shipped two books at +1.000 / −0.750 —MX-asym-TOPandJK-asym-TOP, the second never named.check()now verifies the kind label rather than trusting it: aclipthat does not clip fails too.This corrects this repo's own
CLIPPING_ARM_CORRECTION, whose blast-radius sentence — "campaignB_stats hard-codes its own list and never called check()" — was true in every word and used as an exemption when it was the opposite: not drawing fromcandidates()is exactly why that campaign did not inherit the fix.Also landed, both as negative results
METRIC_DISAGREEMENTin its sharpest form yet. Its "bit-exact" corollary was measured ontorch.randn, where the stated exception cannot occur; on real checkpoints it fails on 0.1–0.9 % of blocks.🤖 Generated with Claude Code