`sprtExpectedN` answers NaN for an error rate at or above 1 and Infinity
for one of 0. Both of its callers gate on
`Number.isFinite(raw) && raw > 0` and quote `maxTrials` when that fails,
and both answers fail it, so `makePlan` and `computeFrontier` returned a
complete plan with every case at the cap and no error anywhere. Neither
entry point validated `alpha` or `beta`, though both validated `mde`,
`fdr` and `pairCoupling`.
`sprtDecision` had the same gap with a louder symptom: an alpha at or
above 1 makes `log(beta / (1 - alpha))` the log of a negative number, so
`lower` came back NaN and the test reported CONTINUE for ever against a
wall that is not a number. `sampleSizeTwoProportion` at an alpha of 1
sets its own critical value to 0 and answers a run count for a design
with no error control in it.
Separately, `sprtExpectedN` returned NaN for `alpha + beta === 1`, which
is inside the domain: both walls sit at a log likelihood ratio of 0 and
Wald's operating characteristic is a genuine 0/0. The answer is 0 runs,
which is what the neighbouring configurations converge to.
No statistics moved, so no README or `paper.test.ts` figure moved.
What broke
sprtExpectedNis the SPRT costing primitive the whole budget layer runs on. It never checked its two error rates, and outside(0, 1)it does not return a wrong number, it returns a non-finite one:NaNfor an alpha at or above 1, becauselog(beta / (1 - alpha))is the log of a negative number, andInfinityfor an alpha of 0.Both of its callers read that the same way:
NaNandInfinityboth fail that test, so both fall back to the per-case cap. NeithermakePlannorevaluatePointvalidatedalphaorbeta, though both validatemde,fdr,pairCoupling,cases,baselineRunsandmaxTrials. The result is a complete plan, with no error anywhere, in which every case is priced atmaxTrials. Measured onmainwithmakePlanover three cases at 51/60, 55/60 and 48/60 and{ ...DEFAULT_PLAN, alpha }:Every one of those returned a
Planobject. Three times the expected bill, and a reader has no way to tell it apart from a suite that genuinely needs the cap. That is the same symptom as #26 ("sprtExpectedNreturned negative run counts ... so affected cases were priced at the per-case cap") from a second cause.Two siblings had the same gap:
sprtDecisionputloweratNaNfor an alpha at or above 1. Every comparison against aNaNis false, so the test reportedCONTINUEfor ever against a wall that is not a number. An alpha of 0 putsupperatInfinity, which no evidence ever crosses.sampleSizeTwoProportionat an alpha of 1 setszAtonormalQuantile(0.5)= 0, which drops the type I error term out of the formula.sampleSizeTwoProportion(0.8, 0.15, 1)answered 29 runs for a design with no error control in it. Its own doc comment calls this "the 'naive approach' peeksafe is measured against", so a plausible number here understates what the naive design costs.And one defect that the guards do not cover, found while writing them:
sprtExpectedNreturnedNaNforalpha + beta === 1, which is inside the domain. Both walls then sit at a log likelihood ratio of 0, so Wald's operating characteristic(1 - B^h) / (A^h - B^h)is a genuine 0/0. Only the pairs where both logs round to exactly 0 in float64 were hit, which is why no guard on the arguments would have caught it:(0.5, 0.5),(0.25, 0.75)and(0.125, 0.875)returnedNaNwhile(0.05, 0.95)escaped on a rounding residual.Why it is a defect and not a judgement call
src/stats.ts's own module header states the contract these five functions broke:brain/architecture/overview.mdrecords that this is the class most merged fixes have been, naming #21, #23, #25 and #27, andsrc/errors.tsalready carries the guard these needed.makePlanandevaluatePointdisagreeing on the same options object is the driftrequirePositiveConfig's own doc comment exists to describe.The fix
Five guards and one closed form, no new helper:
requireOpenProbabilityonalphaandbetainsprtDecision,sprtExpectedNandsampleSizeTwoProportion(src/stats.ts), and oncfg.alphaandcfg.betainmakePlan(src/plan.ts) andevaluatePoint(src/frontier.ts), alongside themdeandfdrchecks already there. Open at both ends: an error rate of 0 is a test that never decides and one of 1 is a test that always does.if (A === B) return 0insprtExpectedN, before the tilt is computed.A === Bisalpha + beta === 1, where the general form(L1 * (A - B) + B) / driftis 0 whateverL1is. The neighbouring configurations already converge to 0 from both sides, which the test pins.A 0 still trips the callers'
Number.isFinite(raw) && raw > 0test and falls back to the cap, which is the right conservative answer for a test that decides before it has seen any data. It now does so on a number rather than on aNaN.No statistic changed for any in-domain input, so no
paper.test.tsnumber and no README figure moved. No runtime dependency, no relaxed refusal, nothing published or tagged.What proves it cannot break the same way
Four new cases, three files. All five assertion groups fail on
mainand pass here:test/budget.test.ts,refuses an out-of-range alpha or beta rather than quoting every case at the cap:makePlanmust throw on 0, 1, -1, 2,NaNandInfinityfor each. It also asserts that the default config prices these cases strictly undercases * maxTrials, so the cap fallback is pinned as a materially different answer and the test cannot be satisfied by a plan that always quotes the cap.test/budget.test.ts,refuses an out-of-range alpha or beta, the same as makePlan: the same six values throughcomputeFrontier.test/stats.test.ts,refuses an error rate outside (0,1) rather than deciding against a NaN wall: the same six throughsprtDecisionandsprtExpectedN, plus assertions that the defaults still give finite walls and a positive expected sample number, so the guard cannot have been satisfied by tightening the domain.test/stats.test.ts,sampleSizeTwoProportion refuses an error rate outside (0,1), which also pins that an effect the case cannot suffer is stillInfinityrather than a refusal: that is an answer about the design, not about the arguments.alpha + beta === 1half is pinned by the three pairs that returnedNaN, each now 0, plus two neighbours that must come in under 1e-3 so the limit is approached and not jumped to.On
mainthese reportexpected [Function] to throw an errorandexpected NaN to be +0.brain/architecture/planner.mdgains two lines under "Things that have been wrong before", next to the #26 entry with the same symptom.Checks
Run on this runner, all green:
npm run typechecknpm test(12 files, 240 tests, up from 236 onmain)npm run buildNot run here and left for CI: Node 22, since this runner is Node 24.21.0, and the packed-tarball dependency check in
.github/workflows/ci.yml. There is no Docker daemon on this runner, though nothing in this repository's checks needs one.Not in flight, and what it overlaps
I read the whole open pull request queue with
gh api repos/sferarc/peeksafe/pulls --paginateand no row limit: 5 rows, #37release/0.1.0, #39chore/pnpm-node26, #40chore/biome-lefthook, #42fix/plan-affordability-grid-baseline, #44ci/hq-stack-check. A scoped search forsprtreturns only #42, and foralphareturns nothing. There are no openhq-queueissues and no issue describes this defect, so there is none to reference; the files changed are named above for the next shift's search.Two overlaps a person resolving should know about:
src/plan.tsandtest/budget.test.tstoo. Different defect and different lines: its guard is inaffordabilityGridand its test is about a missing baseline, mine are inmakePlanand about the error rates. Both are additions to the same two files, so whichever lands second needs the other's lines kept rather than replaced.CHANGELOG.mdentry, deliberately. chore(release): prepare 0.1.0 #37 is rewriting that file to prepare 0.1.0 and an append here would collide with it. fix(plan): the affordability grid priced a suite with no baseline #42 took the same decision. Whoever lands second should add one line under### Fixed.Blocked reviews I could not answer, which outranked this
The shift brief puts a blocked review above anything I pick myself, and
gh pr list --label hq-changes-requestedreturns two: #39 and #40 (draft, stacked on #39). I could not answer either. Every blocking point in the #39 review lands in.github/workflows/release.yml:$HOME/.npmrccredential file,pnpm viewmaking the "version already published" gate one that cannot fail.My charter forbids pushing to
.github/workflows/, and this token holds noworkflowspermission, so GitHub would refuse the push regardless. A doc-only commit would have dropped thehq-changes-requestedlabel and sent #39 back for review with all three blockers still in place, which is worse than leaving it labelled, so I left both untouched. There is a request for a human in my closing message. Nothing else in the queue carries that label.Follow-ups I found and did not take
Three more exported primitives in
src/stats.tsfail the same module-header contract on an optional tuning argument. All are hardening rather than reachable budget errors, since every internal caller passes a valid value, which is why I left them out rather than widening this diff. Measured on this branch:wilsonInterval(5, 10, -2){ low: 0.767, high: 0.233 }, an inverted intervalwilsonInterval(5, 10, NaN){ low: NaN, high: NaN }betaCredibleInterval(betaPosterior(5, 5), 2){ low: 0, high: 1 }, the whole line at 200% credibilitybetaCredibleInterval(betaPosterior(5, 5), -1){ low: 1, high: 0 }, invertedprobabilityMoved(51, 60, 20, 30, -0.5)1, the probability of a drop of minus 50 pointsprobabilityMoved(51, 60, 20, 30, 0.15, 2.5)1, from a Simpson's rule with 3.5 nodessrc/cluster.tsalready has the guard the second row wants, as its ownrequireLevel, with a hint saying "a two-sided confidence level such as 0.95, not a percentage and not an error budget".pairedDiscordancealready has the guard the third row wants on its ownmde, andprobabilityMovedis the last exported entry point taking anmdewithout one. Azneeds a guard no helper covers yet: finite and at or above 0, since azof 0 is a coherent degenerate interval.