Agent-friendly testing, phase 1: rules, guards, one command, goldens, Swift 6 census - #1884
Merged
Merged
Conversation
Makes the test suite friendlier to agents: tests check promises through inputs and outputs, and the checks that go red for no reason stop growing. Rules and guards - Tests/README.md "Test rules", mirrored in AGENTS.md, AGENT_START.md, CLAUDE.md, the agent contract's tests area, the /tests command, and a new test-writer agent that writes tests from the promise, not the code. - scripts/dev/check-test-shape.py: ratchet that blocks NEW tests reading Sources/ as text or asserting on wall-clock time. Existing ones are grandfathered in .agents/test-shape-baseline.json, which only shrinks. - Tests/quarantine.txt: bench a flaky fast-test suite the same day; run-tests.sh skips and lists it, the guard validates the file. - scripts/dev/explain-missing-sources.py: when a hand-kept source list is missing a file, run-tests.sh and the Parakeet lifecycle smoke now print the file that declares the missing name and the list to add it to (replays the Sept 27 integration-smoke failure correctly). - scripts/dev/check-known-traps.py: every Tools package runs in CI and has a matrix rule; every root command is documented. - run-tests.sh defaults TZ to America/Chicago when unset. One command - check.sh: the mapped checks for your diff (agent-check.py), or the quick / full / hardware tiers, with plain PASS/FAIL and rerun commands. Output and real-world layers - MeetingMarkdownGoldenTests + Tests/Fixtures/golden-meetings: whole-file approved copies of saved meeting Markdown, time-zone proof. - scripts/ops/nightly-hardware-smokes.sh: runs check.sh hardware plus the synthetic audio pass on a schedule (not installed; see the doc). - scripts/dev/concurrency-census.sh: typecheck-only strict-concurrency census; 121 warnings across 8 Sources folders, baseline only shrinks. Tests converted from wall-clock limits to outcome checks - TranscriptedConstants detached timeout, ModelLoadProgressWaiter, DefaultInputDeviceMonitor submit, llama child shutdown, and the recovery-owner deadline (counts joins instead of timing the call). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds docs/testing-source-text-inventory.md: all 51 root fast-test files that read Sources/*.swift as text, the promise each one guards, and a verdict (converted, deleted, needs-seam with the smallest seam, or keep). Tests only, no Sources changes. Across 8 files, 12 source-text pins are gone and 19 behavior assertions are in: - AgentConnectionGuide: relocate the capture library between two reads of folderPathsText instead of grepping for `static var`. - AuditRegressionCoverageContract: run the real paster with a non-frontmost target and with a failing dispatcher. - ObservabilityLogWriter: prepare a pre-existing 0644 log through ObservabilityLogFilePreparation and check it comes back 0600 intact. - FocusOrderContract: check settings page identifiers and the ⌘1-⌘5 order through TranscriptedSettingsPage. - StatusItemPresentation: render each MenuBarGlyph and check it is a template drawn in neutral ink. - FailedMeetingPresentation: call retryDisabled for complete, missing, non-retryable, and silent audio. - ClipboardRestoringTextPaster, OverlayScreenSharePrivacy: drop pins already covered by existing behavior suites. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ps linux-checks already runs main's #1874 added two source-text uses to DictationInputDeviceSelectionPolicyTests before the guard existed, so the baseline takes them as grandfathered. #1821 made repo-hygiene run linux-checks.sh, which already includes the new checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…allowlist test The concurrency census flagged Sources/UI going from 53 to 54 warnings after main's Today screen landed. Fixed all three Today warnings instead of grandfathering: the two Int limits on TodayViewModel are nonisolated constants, and the preview's arrowButton helper is @mainactor. UI is now 51, total 119, and the baseline is lowered to match. SentryEventPolicy.allowedDiagnosticTagKeys is internal instead of private, so "Every allowlisted Sentry tag key survives the sanitizer" iterates the real set instead of parsing SentryEventPolicy.swift as text. Also: --help for the nightly and census scripts stops at the header comment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflicts: AgentConnectionGuideTests and FocusOrderContractTests, where the source-text checks had become behavior checks here and main extended the old text checks for the Writing page. Kept the behavior checks and added main's Writing expectations to them (the writing folder follows a relocated library; six pages in ⌘1–⌘6 order). main's Writing PR added 13 source-text reads to UIAutomationSurfaceContractTests before the guard existed; grandfathered. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scripts/dev/mutation-probe.py breaks one Swift file under Sources/ on purpose, one small change at a time (== / !=, < / <=, > / >=, && / ||, true / false, return-bool, single-digit n -> n+1, one-line if negation), runs a test command, and reports KILLED / SURVIVED / COMPILE-ERROR with a mutation score and a JSON report under build/mutation/. It is careful with the real file: refuses paths outside Sources/, untracked or dirty files, and bytes that are not the git index copy; needs a green baseline; restores the original after every mutant, on Ctrl-C/SIGTERM/ SIGHUP and at exit; keeps a backup under build/mutation/.backup; and takes a per-checkout lock. Stdlib only. --list plans without running, --lines re-checks survivors against a wider test command, and --self-test runs offline (wired into linux-checks.sh SELF_TEST_SCRIPTS and the test matrix). First real runs (seed 1, --max 20, each file's own fast test): DictationReadinessWaitPolicy 17 killed / 3 survived (85%), DictationInputDeviceSelectionPolicy 11 / 9 (55%). 8 of those 9 survivors also survive the whole fast suite, including a flipped guard in DictationInputDeviceBindingPolicy.requireSelection. Details and a monthly cadence are in docs/mutation-testing.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d past - a real mic selection passes requireSelection; only device ID 0 is rejected - a Bluetooth-transport headset with a plain name (WH-1000XM5) counts as Bluetooth - with the lid closed, an external mic the user chose is still the one used - ready-start hard recoveries stop exactly when the recovery budget is spent - the readiness refresh gives up at exactly the timeout (exact binary values) Each test was checked red against the one-operator mutation it guards. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same bug as #1883, in the other VM test: linux-checks.sh points TMPDIR deep inside the checkout, which put vnc.py's self-test socket past the 104-byte macOS cap, so 'serve never opened its socket' failed every linux-checks run in a worktree while the standalone run passed. Use a short /tmp directory when the temp root is long. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Last week (Sep 24–28, 200 Swift CI runs) 10 runs went red and none of them was a bug a user would have hit: 6 were a release check tripping on half-finished release PRs, 2 were a wall-clock test on a busy runner, 1 was a hand-kept source list missing a new file, and 1 was a runner glitch. Meanwhile about 1 in 5 fast-test files reads app code as text, and agents spent roughly one test edit for every two app edits keeping tests in sync.
This PR is phase 1 of making the suite agent-friendly: tests check promises through inputs and outputs, and the kinds of checks that go red for no reason stop growing. Plan and before/after: the "Transcripted Testing Plan" doc.
What's in it
Rules (
Tests/README.md"Test rules", mirrored inAGENTS.md,AGENT_START.md,CLAUDE.md, the agent contract's tests area, the/testscommand, and a newtest-writeragent that writes tests from the promise, not the code).Guards
scripts/dev/check-test-shape.py: blocks new tests that readSources/as text or assert on wall-clock time. Existing ones are grandfathered in.agents/test-shape-baseline.json, which can only shrink. Runs inlinux-checks.sh(so repo-hygiene CI) and the test matrix. It caught Skip the engine warmup when the pinned recorder will record #1874 adding two new source-text reads an hour after it was written (grandfathered here since that's on main).Tests/quarantine.txt: bench a flaky fast-test suite the same day;run-tests.shskips and lists it; the guard validates the file.scripts/dev/explain-missing-sources.py: "cannot find 'X' in scope" inrun-tests.shor the Parakeet lifecycle smoke now prints the file that declaresXand the list to add it to. It replays the Sep 27 integration-smoke failure correctly.scripts/dev/check-known-traps.py: every Tools package runs in CI and has a matrix rule; every root command is documented.run-tests.shdefaultsTZto America/Chicago when unset.One command:
bash check.shruns the mapped checks for your diff (viaagent-check.py);quick,fullandhardwaretiers; plain PASS/FAIL with the rerun command.Output and real-world layers
MeetingMarkdownGoldenTests+Tests/Fixtures/golden-meetings/: whole-file approved copies of saved meeting Markdown for three canned meetings. Time-zone proof (passes in Chicago, UTC, Kolkata). Regenerate withTRANSCRIPTED_UPDATE_GOLDENS=1.scripts/ops/nightly-hardware-smokes.sh: runscheck.sh hardwareplus the synthetic audio pass with dated logs and a notification on failure. Not installed: it needs a first manual run for permissions and plays a test tone.scripts/dev/concurrency-census.sh: typecheck-only pass with-strict-concurrency=complete(no binary). 121 concurrency warnings across 8Sources/folders; baseline in.agents/concurrency-baseline.jsoncan only shrink. Mapped toSources/**/*.swiftchanges andcheck.sh full, not added to hosted CI (about 90 s).Tests converted
docs/testing-source-text-inventory.mdlists all 51 files with the promise each guards and the seam each needs.Mutation probe (
scripts/dev/mutation-probe.py,docs/mutation-testing.md): breaks one operator at a time in a Swift file, runs the tests, and lists the breaks nobody catches. It always restores the file (refuses to start on a dirty file, restores on Ctrl-C) and keeps one probe per checkout. First real runs, 20 mutants each:DictationReadinessWaitPolicy.swift: 17 caught, 3 missed (85%).DictationInputDeviceSelectionPolicy.swift: 11 caught, 9 missed (55%). 8 of the 9 also survive the full fast suite. The worst: flippingid != 0inrequireSelectionmakes every real mic selection throw and all fast tests stay green. Also uncaught: a Bluetooth headset with a plain product name, and a user-chosen mic with the lid closed. The newtest-writeragent then wrote tests for all 5 gaps from the promises alone (not the code), and showed each new test goes red under the exact mutation that had slipped through: a real mic passes the selection check, a Bluetooth transport counts even with a plain product name, a user-chosen external mic wins with the lid closed, hard recoveries stop exactly at the budget, and the refresh gives up at exactly its timeout.First seam and Swift 6 fixes (the only production changes)
SentryEventPolicy.allowedDiagnosticTagKeysis internal instead of private, so its test iterates the real set instead of parsing the file.Intlimits onTodayViewModelarenonisolated static let, and the preview'sarrowButtonhelper is@MainActor. 121 to 119.Not in this PR (phase 2 and decisions)
AVAudioEnginecalls inParakeetEngine, movingFailedMeetingItemout ofMeetingSessionController.swift).swift-formatfor new files (it would restyle about 5% of existing lines, mostlyguard ... elsewrapping); makerelease-watch.py"worse" block the next release; keepSpeakerEvalHarnessout of CI.Verification
bash check.shon the branch after merging main at #1788 (40 mapped checks): deterministic proof PASS, every check green.build.sh --no-open,run-tests.sh(19,650 passed),build-deps.sh --force,run-integration-smoke.sh,swift test(588, 1 skipped), Tools package tests,linux-checks.sh(55/55),check-source-pins.py --changed-only,check-test-shape.py,concurrency-census.sh --check(119/119),check-known-traps.py, every new self-test, telemetry-key and analytics-emitter checks.TZ=UTC,America/ChicagoandAsia/Kolkata; a one-word drift fails with the exact line.🤖 Generated with Claude Code