Skip to content

audit(statistics): отзыв заявления о значимости ~10σ; исправление счётчиков и целей - #738

Merged
gHashTag merged 5 commits into
mainfrom
audit/statistics-corrections-2026-08-13
Aug 14, 2026
Merged

gHashTag merged 5 commits into
mainfrom
audit/statistics-corrections-2026-08-13

Conversation

@gHashTag

Copy link
Copy Markdown
Owner

Машинный аудит раздела Statistical Significance в sacred-formulas.md. Девять исправлений, каждое опирается на пересчитанный эталон, а не на процитированное число. Инструмент — каскад «золотое сито» (gHashTag/claim-audit-lab, tools/goldsieve), вердикты воспроизводимы: python3 -m goldsieve run cases/sacred_random_probability.py.

Главное: заявление о значимости ~10σ отзывается

Две вероятности случайного попадания в таблице значимости не вычислялись — они были приняты. Перечисление семейства (V = n\cdot3^k\cdot\pi^m\cdot\varphi^p\cdot e^q) вокруг каждой из 67 фактических целей документа и измерение локальной плотности членов даёт вероятность того, что хотя бы один член попадёт в полосу ±0,01 % вокруг цели:

Перебор Комбинаций Заявлено Пересчитано Занижено
Стандартный ((m \in [-3,0])) 54 756 0,2 % 64,2 % (мин. 33,2 %) ~321×
Расширенный 123 201 2,9 % 99,8 % (мин. 88,4 %) ~34×

При вероятности около 100 % на цель ожидается 67 попаданий из 67 в класс EXACT по случайности. Документ перечисляет 32 — то есть НИЖЕ случайного ожидания. Наблюдение ниже ожидания не подтверждает значимость ни в одну сторону, поэтому «~10σ» и «3σ+» отозваны.

Отдельно: корректный порог значимости выводится из размера перебора. По поправке Шидака при (M = 123,201) и (\alpha = 0{,}05) требуется 5,06σ, а не 3σ.

Совпадения не объявляются неверными — показано, что они неинформативны при точности 0,01 %. Чтобы нести информацию, согласие в этом семействе должно быть точнее локального порога разрешающей способности, а он для этих целей лежит в диапазоне (6{,}6\times10^{-5}\ldots1{,}0\times10^{-4}).

Остальные исправления

Что Было Стало Как проверено
Заголовок Established Constants 75 fits 70 машинный подсчёт строк данных + независимая сумма по подразделам
Класс EXACT 35 32 каждая формула пересчитана из пятёрки параметров; напечатанной колонке Error не доверяем
Время жизни нейтрона «Measured: 879,4 s» 878,4 ± 0,5 с (PDG 2024), +2,0σ сверка с первоисточником
(m_p/m_e) цель 1836,15, ошибка 0,109 % цель 1836,152673426(32), промах 6,3·10⁷σ CODATA 2022; в единицах погрешности измерения, а не в процентах
(T_c) «Unmeasured, 0,0008 % error» измерено, +0,0008σ, но 126 ожидаемых попаданий по случайности HotQCD 2019
(z_{re}) «Testable, 0,005 % error» измерено, 0,0σ, но 1260 ожидаемых попаданий Planck
Размер стандартного перебора 20 412 помечена несогласованность: объявленные границы дают 54 756 перечисление по объявленным границам

Расширенный счёт 123 201 своим границам соответствует и подтверждён.

Достаточность машинной арифметики проверена, а не предположена: ошибка double (8{,}4\times10^{-16}) против допустимых (1{,}7\times10^{-13}), запас 100×. Значит расхождение по (m_p/m_e) не артефакт округления.

Что осталось открытым

  • Поправка Шидака предполагает независимость членов семейства, а они зависимы (общие множители), поэтому реальный порог, вероятно, ниже 5,06σ. Величина зависимости не оценена.
  • Ширина окна усреднения локальной плотности (одна декада) обоснована только устойчивостью результата, теоретического вывода нет.
  • Несогласованность 20 412 против 54 756 помечена, но не устранена: нужно решить, что именно верно — число или границы.

Самокритика

Первая версия проверки сравнивала заявленные 0,2 % с эталоном РАСШИРЕННОГО перебора, тогда как 0,2 % относятся к стандартному. Отсюда в предыдущих отчётах фигурировал множитель «в 500 раз» — он получен на неоднородной паре и недействителен. Здесь каждая колонка сверяется с эталоном своего перебора: 321× и 34×. Вывод о занижении сохраняется по обеим колонкам.

Trinity Audit Loop added 5 commits August 12, 2026 18:44
…lCDFApprox

Three defects of the same shape: a reference value that no one could
regenerate, and code behind a correct name computing the wrong function.

1. src/sacred/zeta_spacing.zig: wignerCDF returned 1 - e^{-x}(1+x),
   x = 4s^2/pi. Its derivative is (32/pi^2) s^3 e^{-4s^2/pi} -- an s^3
   density, not the GUE surmise. Replaced with the exact closed form
   F(s) = erf(2s/sqrt(pi)) - (4s/pi) e^{-4s^2/pi}, plus two tests
   (derivative reproduces the pdf; GUE not GOE, checked at s = 0.3
   because the two curves cross near s = 1).

2. ksPValue (p ~ 2 e^{-2nD^2}) removed. The surmise approximates the exact
   GUE gap law, so its systematic error is fixed while D_crit ~ 1.36/sqrt(n)
   shrinks: at n = 1e5 rejection is guaranteed by construction. Now reports
   D against D_crit(95%) as an effect size. Callers updated in zeta_cf.zig
   and zeta_commands.zig.

3. src/tools/uart_echo_test.zig: normalCDFApprox dropped the exp(-a^2)
   factor of A&S 7.1.26 and mis-nested the polynomial. F(0) = 0.3362 instead
   of 0.5, |error| <= 0.164, range clamped to [0.086, 0.914]. Since the KS
   critical value is 1.36/sqrt(n), the reference error alone exceeded it for
   n > ~250 -- the Normal and Log-Normal latency fits reported 'does not fit'
   for every realistic sample, independently of the data. Fixed and tested.

Documents: data/zeta/zeta_gue_analysis_results.md and
zeta_bin_analysis_update.md carried a GUE reference column of 0.91 / 2.15 /
2.75 for median / p95 / p99 with no derivation. None of the three is
reproducible from the Wigner surmise (0.9639 / 1.7518 / 2.1107), the GOE
surmise, or the s^3 variant above. Recomputed from the 100K Odlyzko zeros in
this repo: p95 deviation is -1.9%, not -19.8%; the 'persistent light tails'
conclusion and the Odlyzko-1989 / Forrester-Mays-2015 confirmations attached
to it are withdrawn. Surviving deviation: std 0.4009 vs 0.4220 (-5.0%).

Reference columns are now computable, not cited:
  scripts/gue_surmise_reference.py       (7 self-tests)
  scripts/recompute_zeta_percentiles.py  (3 self-tests, incl. a Poisson control)

Not compiled: no Zig toolchain in this environment. Zig changes are reviewed
by hand and their numerics verified in Python; run 'zig build test' before merge.
zeta_bin_analysis_update.md read K = 2.6201 +/- 0.0293 vs 2.685 as possible
arithmetic structure in zeta spacings. Monte-Carlo control over uniform random
reals (Khinchin-generic almost surely), same 500 expansions per bin:

  mean of per-expansion geometric means: biased ABOVE K (2.755 at m=20)
  pooled geometric mean over all terms:  biased BELOW K (2.668 at m=20)

The observed value is 1.6 sigma from one control and 4 sigma from the other.
The document records neither the estimator nor the number of partial quotients
per expansion, so the claim is not regenerable either way. Marked OPEN.

scripts/khinchin_finite_sample.py (2 self-tests: recovers K on generic reals,
returns 1 for the golden ratio).
…cts it

A `test` declaration nested inside a struct is not collected by
`zig test` unless that struct is referenced, so the previous in-struct
placement compiled to zero tests: "All 0 tests passed". Verified with
zig 0.15.1 — 1/1 test now runs and passes.
… finite height OPEN

- scripts/gue_exact_gap.py: E2(s) = det(I - K_s) by Nystrom on Gauss-Legendre
  nodes (|dE2| = 4e-16 between n=60 and n=140); the Wigner surmise itself is
  off by 0.3-0.5%, so it cannot serve as the reference for 2% effects.
- scripts/unfolding_test.py: exact Riemann-von Mangoldt unfolding vs the
  leading-order one agree to 1e-5 -- the deficit is not an unfolding artifact.
- scripts/height_extrapolation.py: 1/L extrapolation brings std, p90 and p99
  within 1.5% of exact GUE but leaves p95 at -2.3%; lever arm is too short,
  status OPEN, needs zeros near 10^12.
- scripts/test_reachability.py: only 5492 of 39814 declared tests (<=13.8%)
  are reachable from build.zig roots through the import graph.
…nd targets

Machine audit of the Statistical Significance section (goldsieve cascade,
gHashTag/claim-audit-lab, tools/goldsieve). Nine corrections, each backed by a
recomputed reference rather than a quoted number.

The central problem: the two random-hit probabilities in the significance table
were assumed, not computed. Enumerating the family around each of the 67 actual
targets gives 64.2% for the standard search (claimed 0.2%, understated ~321x)
and 99.8% for the extended search (claimed 2.9%, understated ~34x). At ~100%
per-target probability, 67 of 67 targets are expected in the EXACT class by
chance; the document lists 32 — below chance expectation. The "~10σ" and "3σ+"
statements are therefore withdrawn. A fair threshold at M = 123,201 trials is
5.06σ by the Šidák correction, not 3σ.

Other fixes:
- Established Constants heading 75 -> 70 (machine row count).
- EXACT class count 35 -> 32 (each formula recomputed from its parameter tuple,
  not trusting the printed Error column).
- Neutron lifetime: "Measured: 879.4 s" -> 878.4 ± 0.5 s (PDG 2024), +2.0σ.
- m_p/m_e: target given to full CODATA 2022 precision; 0.109% restated as
  6.3e7σ in units of the measurement uncertainty. Double precision verified
  sufficient (8.4e-16 against an allowance of 1.7e-13).
- T_c and z_re: exact agreement retained but annotated with the number of hits
  expected by chance at that precision (126 and 1260) — agreement inside a
  family this dense carries no information.
- Standard-search size: the stated bounds enumerate 54,756, not 20,412; flagged
  as an inconsistency. The extended count of 123,201 does match its bounds.

The fits are not claimed to be wrong — they are shown to be uninformative at
0.01% precision. The local resolution threshold for these targets is
6.6e-5 … 1.0e-4.
gHashTag pushed a commit to gHashTag/claim-audit-lab that referenced this pull request Aug 13, 2026
Отозван множитель «в 500 раз». Проверка вероятности сравнивала заявленные 0,2 %
(колонка СТАНДАРТНОГО перебора) с эталоном, посчитанным для РАСШИРЕННОГО
перебора 123 201. Знак расхождения верен, величина — нет. Каждая колонка теперь
сверяется с эталоном своего перебора: 321x для стандартного (0,002 против 0,642)
и 34x для расширенного (0,029 против 0,998).

Новое опровержение: объявленный размер стандартного перебора 20 412 против
54 756 комбинаций, которые дают объявленные там же границы. Размер перебора
нельзя брать из текста — от него зависит поправка на множественность.

Приоритет цели. Четыре тика подряд выдали ПОДТВЕРЖДЕНО на печатных значениях
вида phi^phi = 2,17846 — это дефект инструмента: cover сортировал файлы по числу
констант, а таблиц значений в корпусе больше всего. Введена priority(): 2 —
внешнее измерение или погрешность, 1 — статистическое заявление, 0 — печатное
значение. Шесть самопроверок, включая подставку «строка с погрешностью подходит
под шаблон печатного значения».

Самопроверка 84 пройдено, 0 провалено; реестр 36 совпало, 0 изменилось.
Черновик статьи по методике — audits/2026-08-13-goldsieve-v6/vak-draft.md.
Исправления корпуса отданы в gHashTag/trinity#738.
@gHashTag
gHashTag merged commit da2916c into main Aug 14, 2026
17 of 29 checks passed
github-actions Bot added a commit that referenced this pull request Aug 14, 2026
audit(statistics): отзыв заявления о значимости ~10σ; исправление счётчиков и целей (#738)

* fix(zeta): computed GUE reference column; correct wignerCDF and normalCDFApprox

Three defects of the same shape: a reference value that no one could
regenerate, and code behind a correct name computing the wrong function.

1. src/sacred/zeta_spacing.zig: wignerCDF returned 1 - e^{-x}(1+x),
   x = 4s^2/pi. Its derivative is (32/pi^2) s^3 e^{-4s^2/pi} -- an s^3
   density, not the GUE surmise. Replaced with the exact closed form
   F(s) = erf(2s/sqrt(pi)) - (4s/pi) e^{-4s^2/pi}, plus two tests
   (derivative reproduces the pdf; GUE not GOE, checked at s = 0.3
   because the two curves cross near s = 1).

2. ksPValue (p ~ 2 e^{-2nD^2}) removed. The surmise approximates the exact
   GUE gap law, so its systematic error is fixed while D_crit ~ 1.36/sqrt(n)
   shrinks: at n = 1e5 rejection is guaranteed by construction. Now reports
   D against D_crit(95%) as an effect size. Callers updated in zeta_cf.zig
   and zeta_commands.zig.

3. src/tools/uart_echo_test.zig: normalCDFApprox dropped the exp(-a^2)
   factor of A&S 7.1.26 and mis-nested the polynomial. F(0) = 0.3362 instead
   of 0.5, |error| <= 0.164, range clamped to [0.086, 0.914]. Since the KS
   critical value is 1.36/sqrt(n), the reference error alone exceeded it for
   n > ~250 -- the Normal and Log-Normal latency fits reported 'does not fit'
   for every realistic sample, independently of the data. Fixed and tested.

Documents: data/zeta/zeta_gue_analysis_results.md and
zeta_bin_analysis_update.md carried a GUE reference column of 0.91 / 2.15 /
2.75 for median / p95 / p99 with no derivation. None of the three is
reproducible from the Wigner surmise (0.9639 / 1.7518 / 2.1107), the GOE
surmise, or the s^3 variant above. Recomputed from the 100K Odlyzko zeros in
this repo: p95 deviation is -1.9%, not -19.8%; the 'persistent light tails'
conclusion and the Odlyzko-1989 / Forrester-Mays-2015 confirmations attached
to it are withdrawn. Surviving deviation: std 0.4009 vs 0.4220 (-5.0%).

Reference columns are now computable, not cited:
  scripts/gue_surmise_reference.py       (7 self-tests)
  scripts/recompute_zeta_percentiles.py  (3 self-tests, incl. a Poisson control)

Not compiled: no Zig toolchain in this environment. Zig changes are reviewed
by hand and their numerics verified in Python; run 'zig build test' before merge.

* audit(zeta): Khinchin K deficit is estimator-dependent, not a finding

zeta_bin_analysis_update.md read K = 2.6201 +/- 0.0293 vs 2.685 as possible
arithmetic structure in zeta spacings. Monte-Carlo control over uniform random
reals (Khinchin-generic almost surely), same 500 expansions per bin:

  mean of per-expansion geometric means: biased ABOVE K (2.755 at m=20)
  pooled geometric mean over all terms:  biased BELOW K (2.668 at m=20)

The observed value is 1.6 sigma from one control and 4 sigma from the other.
The document records neither the estimator nor the number of partial quotients
per expansion, so the claim is not regenerable either way. Marked OPEN.

scripts/khinchin_finite_sample.py (2 self-tests: recovers K on generic reals,
returns 1 for the golden ratio).

* test(uart): move normalCDFApprox test to file scope so zig test collects it

A `test` declaration nested inside a struct is not collected by
`zig test` unless that struct is referenced, so the previous in-struct
placement compiled to zero tests: "All 0 tests passed". Verified with
zig 0.15.1 — 1/1 test now runs and passes.

* audit(zeta): exact GUE gap law as the reference; unfolding ruled out, finite height OPEN

- scripts/gue_exact_gap.py: E2(s) = det(I - K_s) by Nystrom on Gauss-Legendre
  nodes (|dE2| = 4e-16 between n=60 and n=140); the Wigner surmise itself is
  off by 0.3-0.5%, so it cannot serve as the reference for 2% effects.
- scripts/unfolding_test.py: exact Riemann-von Mangoldt unfolding vs the
  leading-order one agree to 1e-5 -- the deficit is not an unfolding artifact.
- scripts/height_extrapolation.py: 1/L extrapolation brings std, p90 and p99
  within 1.5% of exact GUE but leaves p95 at -2.3%; lever arm is too short,
  status OPEN, needs zeros near 10^12.
- scripts/test_reachability.py: only 5492 of 39814 declared tests (<=13.8%)
  are reachable from build.zig roots through the import graph.

* audit(statistics): withdraw the ~10σ significance claim; fix counts and targets

Machine audit of the Statistical Significance section (goldsieve cascade,
gHashTag/claim-audit-lab, tools/goldsieve). Nine corrections, each backed by a
recomputed reference rather than a quoted number.

The central problem: the two random-hit probabilities in the significance table
were assumed, not computed. Enumerating the family around each of the 67 actual
targets gives 64.2% for the standard search (claimed 0.2%, understated ~321x)
and 99.8% for the extended search (claimed 2.9%, understated ~34x). At ~100%
per-target probability, 67 of 67 targets are expected in the EXACT class by
chance; the document lists 32 — below chance expectation. The "~10σ" and "3σ+"
statements are therefore withdrawn. A fair threshold at M = 123,201 trials is
5.06σ by the Šidák correction, not 3σ.

Other fixes:
- Established Constants heading 75 -> 70 (machine row count).
- EXACT class count 35 -> 32 (each formula recomputed from its parameter tuple,
  not trusting the printed Error column).
- Neutron lifetime: "Measured: 879.4 s" -> 878.4 ± 0.5 s (PDG 2024), +2.0σ.
- m_p/m_e: target given to full CODATA 2022 precision; 0.109% restated as
  6.3e7σ in units of the measurement uncertainty. Double precision verified
  sufficient (8.4e-16 against an allowance of 1.7e-13).
- T_c and z_re: exact agreement retained but annotated with the number of hits
  expected by chance at that precision (126 and 1260) — agreement inside a
  family this dense carries no information.
- Standard-search size: the stated bounds enumerate 54,756, not 20,412; flagged
  as an inconsistency. The extended count of 123,201 does match its bounds.

The fits are not claimed to be wrong — they are shown to be uninformative at
0.01% precision. The local resolution threshold for these targets is
6.6e-5 … 1.0e-4.

---------

Co-authored-by: Trinity Audit Loop <oxicocicate35@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant