Repository navigation
audit(statistics): отзыв заявления о значимости ~10σ; исправление счётчиков и целей - #738
Merged
Merged
Conversation
added 5 commits
August 12, 2026 18:44
…lCDFApprox
Three defects of the same shape: a reference value that no one could
regenerate, and code behind a correct name computing the wrong function.
1. src/sacred/zeta_spacing.zig: wignerCDF returned 1 - e^{-x}(1+x),
x = 4s^2/pi. Its derivative is (32/pi^2) s^3 e^{-4s^2/pi} -- an s^3
density, not the GUE surmise. Replaced with the exact closed form
F(s) = erf(2s/sqrt(pi)) - (4s/pi) e^{-4s^2/pi}, plus two tests
(derivative reproduces the pdf; GUE not GOE, checked at s = 0.3
because the two curves cross near s = 1).
2. ksPValue (p ~ 2 e^{-2nD^2}) removed. The surmise approximates the exact
GUE gap law, so its systematic error is fixed while D_crit ~ 1.36/sqrt(n)
shrinks: at n = 1e5 rejection is guaranteed by construction. Now reports
D against D_crit(95%) as an effect size. Callers updated in zeta_cf.zig
and zeta_commands.zig.
3. src/tools/uart_echo_test.zig: normalCDFApprox dropped the exp(-a^2)
factor of A&S 7.1.26 and mis-nested the polynomial. F(0) = 0.3362 instead
of 0.5, |error| <= 0.164, range clamped to [0.086, 0.914]. Since the KS
critical value is 1.36/sqrt(n), the reference error alone exceeded it for
n > ~250 -- the Normal and Log-Normal latency fits reported 'does not fit'
for every realistic sample, independently of the data. Fixed and tested.
Documents: data/zeta/zeta_gue_analysis_results.md and
zeta_bin_analysis_update.md carried a GUE reference column of 0.91 / 2.15 /
2.75 for median / p95 / p99 with no derivation. None of the three is
reproducible from the Wigner surmise (0.9639 / 1.7518 / 2.1107), the GOE
surmise, or the s^3 variant above. Recomputed from the 100K Odlyzko zeros in
this repo: p95 deviation is -1.9%, not -19.8%; the 'persistent light tails'
conclusion and the Odlyzko-1989 / Forrester-Mays-2015 confirmations attached
to it are withdrawn. Surviving deviation: std 0.4009 vs 0.4220 (-5.0%).
Reference columns are now computable, not cited:
scripts/gue_surmise_reference.py (7 self-tests)
scripts/recompute_zeta_percentiles.py (3 self-tests, incl. a Poisson control)
Not compiled: no Zig toolchain in this environment. Zig changes are reviewed
by hand and their numerics verified in Python; run 'zig build test' before merge.
zeta_bin_analysis_update.md read K = 2.6201 +/- 0.0293 vs 2.685 as possible arithmetic structure in zeta spacings. Monte-Carlo control over uniform random reals (Khinchin-generic almost surely), same 500 expansions per bin: mean of per-expansion geometric means: biased ABOVE K (2.755 at m=20) pooled geometric mean over all terms: biased BELOW K (2.668 at m=20) The observed value is 1.6 sigma from one control and 4 sigma from the other. The document records neither the estimator nor the number of partial quotients per expansion, so the claim is not regenerable either way. Marked OPEN. scripts/khinchin_finite_sample.py (2 self-tests: recovers K on generic reals, returns 1 for the golden ratio).
…cts it A `test` declaration nested inside a struct is not collected by `zig test` unless that struct is referenced, so the previous in-struct placement compiled to zero tests: "All 0 tests passed". Verified with zig 0.15.1 — 1/1 test now runs and passes.
… finite height OPEN - scripts/gue_exact_gap.py: E2(s) = det(I - K_s) by Nystrom on Gauss-Legendre nodes (|dE2| = 4e-16 between n=60 and n=140); the Wigner surmise itself is off by 0.3-0.5%, so it cannot serve as the reference for 2% effects. - scripts/unfolding_test.py: exact Riemann-von Mangoldt unfolding vs the leading-order one agree to 1e-5 -- the deficit is not an unfolding artifact. - scripts/height_extrapolation.py: 1/L extrapolation brings std, p90 and p99 within 1.5% of exact GUE but leaves p95 at -2.3%; lever arm is too short, status OPEN, needs zeros near 10^12. - scripts/test_reachability.py: only 5492 of 39814 declared tests (<=13.8%) are reachable from build.zig roots through the import graph.
…nd targets Machine audit of the Statistical Significance section (goldsieve cascade, gHashTag/claim-audit-lab, tools/goldsieve). Nine corrections, each backed by a recomputed reference rather than a quoted number. The central problem: the two random-hit probabilities in the significance table were assumed, not computed. Enumerating the family around each of the 67 actual targets gives 64.2% for the standard search (claimed 0.2%, understated ~321x) and 99.8% for the extended search (claimed 2.9%, understated ~34x). At ~100% per-target probability, 67 of 67 targets are expected in the EXACT class by chance; the document lists 32 — below chance expectation. The "~10σ" and "3σ+" statements are therefore withdrawn. A fair threshold at M = 123,201 trials is 5.06σ by the Šidák correction, not 3σ. Other fixes: - Established Constants heading 75 -> 70 (machine row count). - EXACT class count 35 -> 32 (each formula recomputed from its parameter tuple, not trusting the printed Error column). - Neutron lifetime: "Measured: 879.4 s" -> 878.4 ± 0.5 s (PDG 2024), +2.0σ. - m_p/m_e: target given to full CODATA 2022 precision; 0.109% restated as 6.3e7σ in units of the measurement uncertainty. Double precision verified sufficient (8.4e-16 against an allowance of 1.7e-13). - T_c and z_re: exact agreement retained but annotated with the number of hits expected by chance at that precision (126 and 1260) — agreement inside a family this dense carries no information. - Standard-search size: the stated bounds enumerate 54,756, not 20,412; flagged as an inconsistency. The extended count of 123,201 does match its bounds. The fits are not claimed to be wrong — they are shown to be uninformative at 0.01% precision. The local resolution threshold for these targets is 6.6e-5 … 1.0e-4.
gHashTag
pushed a commit
to gHashTag/claim-audit-lab
that referenced
this pull request
Aug 13, 2026
Отозван множитель «в 500 раз». Проверка вероятности сравнивала заявленные 0,2 % (колонка СТАНДАРТНОГО перебора) с эталоном, посчитанным для РАСШИРЕННОГО перебора 123 201. Знак расхождения верен, величина — нет. Каждая колонка теперь сверяется с эталоном своего перебора: 321x для стандартного (0,002 против 0,642) и 34x для расширенного (0,029 против 0,998). Новое опровержение: объявленный размер стандартного перебора 20 412 против 54 756 комбинаций, которые дают объявленные там же границы. Размер перебора нельзя брать из текста — от него зависит поправка на множественность. Приоритет цели. Четыре тика подряд выдали ПОДТВЕРЖДЕНО на печатных значениях вида phi^phi = 2,17846 — это дефект инструмента: cover сортировал файлы по числу констант, а таблиц значений в корпусе больше всего. Введена priority(): 2 — внешнее измерение или погрешность, 1 — статистическое заявление, 0 — печатное значение. Шесть самопроверок, включая подставку «строка с погрешностью подходит под шаблон печатного значения». Самопроверка 84 пройдено, 0 провалено; реестр 36 совпало, 0 изменилось. Черновик статьи по методике — audits/2026-08-13-goldsieve-v6/vak-draft.md. Исправления корпуса отданы в gHashTag/trinity#738.
github-actions Bot
added a commit
that referenced
this pull request
Aug 14, 2026
audit(statistics): отзыв заявления о значимости ~10σ; исправление счётчиков и целей (#738) * fix(zeta): computed GUE reference column; correct wignerCDF and normalCDFApprox Three defects of the same shape: a reference value that no one could regenerate, and code behind a correct name computing the wrong function. 1. src/sacred/zeta_spacing.zig: wignerCDF returned 1 - e^{-x}(1+x), x = 4s^2/pi. Its derivative is (32/pi^2) s^3 e^{-4s^2/pi} -- an s^3 density, not the GUE surmise. Replaced with the exact closed form F(s) = erf(2s/sqrt(pi)) - (4s/pi) e^{-4s^2/pi}, plus two tests (derivative reproduces the pdf; GUE not GOE, checked at s = 0.3 because the two curves cross near s = 1). 2. ksPValue (p ~ 2 e^{-2nD^2}) removed. The surmise approximates the exact GUE gap law, so its systematic error is fixed while D_crit ~ 1.36/sqrt(n) shrinks: at n = 1e5 rejection is guaranteed by construction. Now reports D against D_crit(95%) as an effect size. Callers updated in zeta_cf.zig and zeta_commands.zig. 3. src/tools/uart_echo_test.zig: normalCDFApprox dropped the exp(-a^2) factor of A&S 7.1.26 and mis-nested the polynomial. F(0) = 0.3362 instead of 0.5, |error| <= 0.164, range clamped to [0.086, 0.914]. Since the KS critical value is 1.36/sqrt(n), the reference error alone exceeded it for n > ~250 -- the Normal and Log-Normal latency fits reported 'does not fit' for every realistic sample, independently of the data. Fixed and tested. Documents: data/zeta/zeta_gue_analysis_results.md and zeta_bin_analysis_update.md carried a GUE reference column of 0.91 / 2.15 / 2.75 for median / p95 / p99 with no derivation. None of the three is reproducible from the Wigner surmise (0.9639 / 1.7518 / 2.1107), the GOE surmise, or the s^3 variant above. Recomputed from the 100K Odlyzko zeros in this repo: p95 deviation is -1.9%, not -19.8%; the 'persistent light tails' conclusion and the Odlyzko-1989 / Forrester-Mays-2015 confirmations attached to it are withdrawn. Surviving deviation: std 0.4009 vs 0.4220 (-5.0%). Reference columns are now computable, not cited: scripts/gue_surmise_reference.py (7 self-tests) scripts/recompute_zeta_percentiles.py (3 self-tests, incl. a Poisson control) Not compiled: no Zig toolchain in this environment. Zig changes are reviewed by hand and their numerics verified in Python; run 'zig build test' before merge. * audit(zeta): Khinchin K deficit is estimator-dependent, not a finding zeta_bin_analysis_update.md read K = 2.6201 +/- 0.0293 vs 2.685 as possible arithmetic structure in zeta spacings. Monte-Carlo control over uniform random reals (Khinchin-generic almost surely), same 500 expansions per bin: mean of per-expansion geometric means: biased ABOVE K (2.755 at m=20) pooled geometric mean over all terms: biased BELOW K (2.668 at m=20) The observed value is 1.6 sigma from one control and 4 sigma from the other. The document records neither the estimator nor the number of partial quotients per expansion, so the claim is not regenerable either way. Marked OPEN. scripts/khinchin_finite_sample.py (2 self-tests: recovers K on generic reals, returns 1 for the golden ratio). * test(uart): move normalCDFApprox test to file scope so zig test collects it A `test` declaration nested inside a struct is not collected by `zig test` unless that struct is referenced, so the previous in-struct placement compiled to zero tests: "All 0 tests passed". Verified with zig 0.15.1 — 1/1 test now runs and passes. * audit(zeta): exact GUE gap law as the reference; unfolding ruled out, finite height OPEN - scripts/gue_exact_gap.py: E2(s) = det(I - K_s) by Nystrom on Gauss-Legendre nodes (|dE2| = 4e-16 between n=60 and n=140); the Wigner surmise itself is off by 0.3-0.5%, so it cannot serve as the reference for 2% effects. - scripts/unfolding_test.py: exact Riemann-von Mangoldt unfolding vs the leading-order one agree to 1e-5 -- the deficit is not an unfolding artifact. - scripts/height_extrapolation.py: 1/L extrapolation brings std, p90 and p99 within 1.5% of exact GUE but leaves p95 at -2.3%; lever arm is too short, status OPEN, needs zeros near 10^12. - scripts/test_reachability.py: only 5492 of 39814 declared tests (<=13.8%) are reachable from build.zig roots through the import graph. * audit(statistics): withdraw the ~10σ significance claim; fix counts and targets Machine audit of the Statistical Significance section (goldsieve cascade, gHashTag/claim-audit-lab, tools/goldsieve). Nine corrections, each backed by a recomputed reference rather than a quoted number. The central problem: the two random-hit probabilities in the significance table were assumed, not computed. Enumerating the family around each of the 67 actual targets gives 64.2% for the standard search (claimed 0.2%, understated ~321x) and 99.8% for the extended search (claimed 2.9%, understated ~34x). At ~100% per-target probability, 67 of 67 targets are expected in the EXACT class by chance; the document lists 32 — below chance expectation. The "~10σ" and "3σ+" statements are therefore withdrawn. A fair threshold at M = 123,201 trials is 5.06σ by the Šidák correction, not 3σ. Other fixes: - Established Constants heading 75 -> 70 (machine row count). - EXACT class count 35 -> 32 (each formula recomputed from its parameter tuple, not trusting the printed Error column). - Neutron lifetime: "Measured: 879.4 s" -> 878.4 ± 0.5 s (PDG 2024), +2.0σ. - m_p/m_e: target given to full CODATA 2022 precision; 0.109% restated as 6.3e7σ in units of the measurement uncertainty. Double precision verified sufficient (8.4e-16 against an allowance of 1.7e-13). - T_c and z_re: exact agreement retained but annotated with the number of hits expected by chance at that precision (126 and 1260) — agreement inside a family this dense carries no information. - Standard-search size: the stated bounds enumerate 54,756, not 20,412; flagged as an inconsistency. The extended count of 123,201 does match its bounds. The fits are not claimed to be wrong — they are shown to be uninformative at 0.01% precision. The local resolution threshold for these targets is 6.6e-5 … 1.0e-4. --------- Co-authored-by: Trinity Audit Loop <oxicocicate35@gmail.com>
This was referenced Aug 14, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Машинный аудит раздела Statistical Significance в
sacred-formulas.md. Девять исправлений, каждое опирается на пересчитанный эталон, а не на процитированное число. Инструмент — каскад «золотое сито» (gHashTag/claim-audit-lab,tools/goldsieve), вердикты воспроизводимы:python3 -m goldsieve run cases/sacred_random_probability.py.Главное: заявление о значимости ~10σ отзывается
Две вероятности случайного попадания в таблице значимости не вычислялись — они были приняты. Перечисление семейства (V = n\cdot3^k\cdot\pi^m\cdot\varphi^p\cdot e^q) вокруг каждой из 67 фактических целей документа и измерение локальной плотности членов даёт вероятность того, что хотя бы один член попадёт в полосу ±0,01 % вокруг цели:
При вероятности около 100 % на цель ожидается 67 попаданий из 67 в класс EXACT по случайности. Документ перечисляет 32 — то есть НИЖЕ случайного ожидания. Наблюдение ниже ожидания не подтверждает значимость ни в одну сторону, поэтому «~10σ» и «3σ+» отозваны.
Отдельно: корректный порог значимости выводится из размера перебора. По поправке Шидака при (M = 123,201) и (\alpha = 0{,}05) требуется 5,06σ, а не 3σ.
Совпадения не объявляются неверными — показано, что они неинформативны при точности 0,01 %. Чтобы нести информацию, согласие в этом семействе должно быть точнее локального порога разрешающей способности, а он для этих целей лежит в диапазоне (6{,}6\times10^{-5}\ldots1{,}0\times10^{-4}).
Остальные исправления
Расширенный счёт 123 201 своим границам соответствует и подтверждён.
Достаточность машинной арифметики проверена, а не предположена: ошибка double (8{,}4\times10^{-16}) против допустимых (1{,}7\times10^{-13}), запас 100×. Значит расхождение по (m_p/m_e) не артефакт округления.
Что осталось открытым
Самокритика
Первая версия проверки сравнивала заявленные 0,2 % с эталоном РАСШИРЕННОГО перебора, тогда как 0,2 % относятся к стандартному. Отсюда в предыдущих отчётах фигурировал множитель «в 500 раз» — он получен на неоднородной паре и недействителен. Здесь каждая колонка сверяется с эталоном своего перебора: 321× и 34×. Вывод о занижении сохраняется по обеим колонкам.