Skip to content

fix(research): eleven verdicts were windows pretending to be checkpoints - #563

Merged
gHashTag merged 1 commit into
mainfrom
fix/pooled-verdicts-and-assert-sweep
Aug 12, 2026
Merged

gHashTag merged 1 commit into
mainfrom
fix/pooled-verdicts-and-assert-sweep

Conversation

@gHashTag

Copy link
Copy Markdown
Owner

campaignB_stats.row() did this:

d = np.concatenate([dvec(D, m, arm, ref) for m in models])   # 4 models × 35 windows
r = paired(d)                                                # …as n = 140

Windows replicate the text, not the model family. Sixth instance of this error in the campaign, and the first found in a script rather than in prose — the previous five were caught by re-reading documents; this one was still executing. The section header said POOLED OVER ALL FOUR MODELS while the statistic said n = 140 windows.

Eleven of fourteen verdicts flip

arm vs window-pooled n=140 model-level n=4
MX-asym-NEAR0 MXFP4 −4.99 %, p=1.6e-44 BEATS −4.76 % [−8.58, −0.78], p=0.032 TIE
MX-asym-MID MXFP4 −2.21 %, p=9.6e-26 BEATS −2.08 % [−3.48, −0.67], p=0.019 TIE
MX-asym-NEAR0 NF4 −0.92 %, p=1.6e-02 BEATS −0.92 % [−6.62, +5.13], p=0.655 TIE
MX-asym-MID NF4 +1.98 %, p=9.0e-06 loses +1.87 % [−4.68, +8.87], p=0.440 TIE

…and seven more. The point estimates barely move (−4.99 % → −4.76 %); the intervals grow by ≈√35 and every one now contains zero. One verdict survives, and its own tag already says 3/4 in-sample.

The correction is symmetric — four rows moved toward our codebooks, seven away. An error that only ever flattered us would be a different finding.

Consequences, restated in the documents that quoted the pooled figures

  • T40's decomposition: (+0.464 %) × (−4.324 %) = −3.880 %, residual 6.94e-18. The composition is arithmetic and holds at either level; what is withdrawn is the strength of the surrounding margins. At four checkpoints "NF4 beats MXFP4" is a TIE (−3.88 %, p = 0.208).
  • MX-asym-NEAR0's "−4.74 % held-out, 4/4": the rotation is unanimous, so that mean is algebraically identical to the plain four-model mean. The held-out label transports nothing. Rotation-honest counterpart on the full nine-arm pool: −2.46 %, three distinct arms — a protocol result with no codebook to attach it to. "The only arm clearing significance" is false: four clear uncorrected, none clears Bonferroni over the nine the argmin came from.

The assert sweep (187 sites) found the clipping arm had a twin

campaignB_books asserted max(abs(x)) under a docstring reading BOTH tails and shipped two books at +1.000 / −0.750 — MX-asym-TOP and JK-asym-TOP, the second never named. check() now verifies the kind label rather than trusting it: a clip that does not clip fails too.

This corrects this repo's own CLIPPING_ARM_CORRECTION, whose blast-radius sentence — "campaignB_stats hard-codes its own list and never called check()" — was true in every word and used as an exemption when it was the opposite: not drawing from candidates() is exactly why that campaign did not inherit the fix.

Also landed, both as negative results

  • T41 REFUTED, with its literature check (ACIQ, LAPQ, Four Over Six, MXAttention/UOS — most of the trade is published). The arm lowers squared weight error by 6.8–11.3 % on every checkpoint while raising perplexity on two, which is METRIC_DISAGREEMENT in its sharpest form yet. Its "bit-exact" corollary was measured on torch.randn, where the stated exception cannot occur; on real checkpoints it fails on 0.1–0.9 % of blocks.
  • The seed-vs-distance control ran and did not break the confound. All four Lloyd-seeded fits landed far; two converged there rather than running out of budget. Reported as such, with the seed hypothesis's decisive evidence replicating 1 of 4.

🤖 Generated with Claude Code

campaignB_stats.row() did `np.concatenate([dvec(D, m, arm, ref) for m in
models])` and handed paired() 140 windows from four models as 140 replicates.
Windows replicate the TEXT, not the model family. This is the sixth instance of
this error in the campaign and the first found in a script rather than in prose
-- the previous five were caught by re-reading documents; this one was still
executing.

row() now takes n = models for cross-model claims, each contributing its own
mean log-ratio, and keeps windows only for single-model rows, which are
within-model claims entitled to them. Eleven of fourteen verdicts flip: every
"BEATS MXFP4" and every "loses to NF4" becomes a TIE. Point estimates barely
move (-4.99% -> -4.76%); the intervals grow by ~sqrt(35) and all now contain
zero. The correction is symmetric -- four rows moved toward our codebooks, seven
away -- which is how it is distinguishable from a motivated one.

Restated at the model level, in the documents that quoted the pooled figures:
- T40's decomposition: (+0.464%) x (-4.324%) = -3.880%, residual 6.94e-18. The
  decomposition is arithmetic and holds at either level; what is withdrawn is
  the strength of the surrounding margins. At four checkpoints "NF4 beats MXFP4"
  is a TIE (-3.88%, p = 0.208) and NF4's win is a per-model result on 3 of 4.
- MX-asym-NEAR0's "-4.74% held-out, 4/4": the rotation is UNANIMOUS, so that
  mean is algebraically identical to the plain four-model mean. The held-out
  label transports nothing. Rotation-honest counterpart on the full nine-arm
  pool: -2.46%, three distinct arms, so it is a protocol result with no codebook
  to attach it to. "The only arm clearing significance" is false: four clear
  uncorrected, none clears Bonferroni over the nine the argmin came from.

The assert sweep (187 sites) found the clipping-arm defect had an unfixed twin:
campaignB_books asserted max(abs(x)) under a docstring reading "BOTH tails", and
shipped TWO books at +1.000/-0.750 -- MX-asym-TOP and JK-asym-TOP, the second
never named. check() now verifies the kind label rather than trusting it: a
"clip" that does not clip fails too. The Bonferroni family is three placements,
not four; the reclassification was made on structural grounds a day before
anyone computed which way it moved a verdict, and at model level it moves none.

Corrects this repo's own CLIPPING_ARM_CORRECTION, whose blast-radius sentence
("campaignB_stats hard-codes its own list and never called check()") was true in
every word and used as an exemption when it was the opposite: not drawing from
candidates() is exactly why that campaign did not inherit the fix.

T41 is recorded as REFUTED with its literature check: the arm lowers squared
weight error by 6.8-11.3% on every checkpoint while raising perplexity on two,
and its "bit-exact" corollary was measured on torch.randn where the stated
exception cannot occur -- on real checkpoints it fails on 0.1-0.9% of blocks.
The seed-vs-distance control ran and did NOT break the confound; reported as
such, with the seed hypothesis's decisive evidence replicating 1 of 4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gHashTag
gHashTag merged commit f4e361a into main Aug 12, 2026
9 of 13 checks passed
@gHashTag
gHashTag deleted the fix/pooled-verdicts-and-assert-sweep branch August 12, 2026 15:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant