Skip to content

Agent-friendly testing, phase 1: rules, guards, one command, goldens, Swift 6 census - #1884

Merged
r3dbars merged 14 commits into
mainfrom
claude/ai-unit-tests-theory-9ec0d0
Sep 28, 2026
Merged

r3dbars merged 14 commits into
mainfrom
claude/ai-unit-tests-theory-9ec0d0

Conversation

@r3dbars

@r3dbars r3dbars commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Why

Last week (Sep 24–28, 200 Swift CI runs) 10 runs went red and none of them was a bug a user would have hit: 6 were a release check tripping on half-finished release PRs, 2 were a wall-clock test on a busy runner, 1 was a hand-kept source list missing a new file, and 1 was a runner glitch. Meanwhile about 1 in 5 fast-test files reads app code as text, and agents spent roughly one test edit for every two app edits keeping tests in sync.

This PR is phase 1 of making the suite agent-friendly: tests check promises through inputs and outputs, and the kinds of checks that go red for no reason stop growing. Plan and before/after: the "Transcripted Testing Plan" doc.

What's in it

Rules (Tests/README.md "Test rules", mirrored in AGENTS.md, AGENT_START.md, CLAUDE.md, the agent contract's tests area, the /tests command, and a new test-writer agent that writes tests from the promise, not the code).

Guards

  • scripts/dev/check-test-shape.py: blocks new tests that read Sources/ as text or assert on wall-clock time. Existing ones are grandfathered in .agents/test-shape-baseline.json, which can only shrink. Runs in linux-checks.sh (so repo-hygiene CI) and the test matrix. It caught Skip the engine warmup when the pinned recorder will record #1874 adding two new source-text reads an hour after it was written (grandfathered here since that's on main).
  • Tests/quarantine.txt: bench a flaky fast-test suite the same day; run-tests.sh skips and lists it; the guard validates the file.
  • scripts/dev/explain-missing-sources.py: "cannot find 'X' in scope" in run-tests.sh or the Parakeet lifecycle smoke now prints the file that declares X and the list to add it to. It replays the Sep 27 integration-smoke failure correctly.
  • scripts/dev/check-known-traps.py: every Tools package runs in CI and has a matrix rule; every root command is documented.
  • run-tests.sh defaults TZ to America/Chicago when unset.

One command: bash check.sh runs the mapped checks for your diff (via agent-check.py); quick, full and hardware tiers; plain PASS/FAIL with the rerun command.

Output and real-world layers

  • MeetingMarkdownGoldenTests + Tests/Fixtures/golden-meetings/: whole-file approved copies of saved meeting Markdown for three canned meetings. Time-zone proof (passes in Chicago, UTC, Kolkata). Regenerate with TRANSCRIPTED_UPDATE_GOLDENS=1.
  • scripts/ops/nightly-hardware-smokes.sh: runs check.sh hardware plus the synthetic audio pass with dated logs and a notification on failure. Not installed: it needs a first manual run for permissions and plays a test tone.
  • scripts/dev/concurrency-census.sh: typecheck-only pass with -strict-concurrency=complete (no binary). 121 concurrency warnings across 8 Sources/ folders; baseline in .agents/concurrency-baseline.json can only shrink. Mapped to Sources/**/*.swift changes and check.sh full, not added to hosted CI (about 90 s).

Tests converted

  • 5 wall-clock assertions became outcome checks (detached timeout, model-load waiter, default-input submit, llama child shutdown, recovery-owner deadline counts joins instead of timing).
  • 8 source-text files converted fully or partly (12 source reads out, 19 behavior assertions in); docs/testing-source-text-inventory.md lists all 51 files with the promise each guards and the seam each needs.

Mutation probe (scripts/dev/mutation-probe.py, docs/mutation-testing.md): breaks one operator at a time in a Swift file, runs the tests, and lists the breaks nobody catches. It always restores the file (refuses to start on a dirty file, restores on Ctrl-C) and keeps one probe per checkout. First real runs, 20 mutants each:

  • DictationReadinessWaitPolicy.swift: 17 caught, 3 missed (85%).
  • DictationInputDeviceSelectionPolicy.swift: 11 caught, 9 missed (55%). 8 of the 9 also survive the full fast suite. The worst: flipping id != 0 in requireSelection makes every real mic selection throw and all fast tests stay green. Also uncaught: a Bluetooth headset with a plain product name, and a user-chosen mic with the lid closed. The new test-writer agent then wrote tests for all 5 gaps from the promises alone (not the code), and showed each new test goes red under the exact mutation that had slipped through: a real mic passes the selection check, a Bluetooth transport counts even with a plain product name, a user-chosen external mic wins with the lid closed, hard recoveries stop exactly at the budget, and the refresh gives up at exactly its timeout.

First seam and Swift 6 fixes (the only production changes)

  • SentryEventPolicy.allowedDiagnosticTagKeys is internal instead of private, so its test iterates the real set instead of parsing the file.
  • The census flagged a new warning from Today: day sentence, writing lane, and a preview of each capture #1873 (Today). Fixed all three Today warnings: two Int limits on TodayViewModel are nonisolated static let, and the preview's arrowButton helper is @MainActor. 121 to 119.

Not in this PR (phase 2 and decisions)

  • 38 source-text files need a production seam first (top three: the dictation stop/start pipeline, a protocol around AVAudioEngine calls in ParakeetEngine, moving FailedMeetingItem out of MeetingSessionController.swift).
  • Fakes for clock, mic, clipboard and STT; state enums for flag-heavy types; Swift 6 per-folder flips.
  • Decisions: install the nightly runner and at what time; adopt swift-format for new files (it would restyle about 5% of existing lines, mostly guard ... else wrapping); make release-watch.py "worse" block the next release; keep SpeakerEvalHarness out of CI.

Verification

bash check.sh on the branch after merging main at #1788 (40 mapped checks): deterministic proof PASS, every check green.

  • build.sh --no-open, run-tests.sh (19,650 passed), build-deps.sh --force, run-integration-smoke.sh, swift test (588, 1 skipped), Tools package tests, linux-checks.sh (55/55), check-source-pins.py --changed-only, check-test-shape.py, concurrency-census.sh --check (119/119), check-known-traps.py, every new self-test, telemetry-key and analytics-emitter checks.
  • Golden tests pass under TZ=UTC, America/Chicago and Asia/Kolkata; a one-word drift fails with the exact line.
  • Quarantine demo: a benched suite is skipped and listed. Missing-source demo: dropping a file from APP_SOURCES makes the runner name it.
  • Every merge of main so far needed the new guards' baselines re-grandfathered: Skip the engine warmup when the pinned recorder will record #1874, the Writing tab and Make system-audio meeting warnings actionable #1761 added 18 source-text reads between them before the guard existed. After this merges, new ones fail in CI instead.
  • Not run: hardware smokes (need mic and System Audio permission), the nightly installer.

🤖 Generated with Claude Code

r3dbars and others added 14 commits September 27, 2026 21:09
Makes the test suite friendlier to agents: tests check promises through
inputs and outputs, and the checks that go red for no reason stop growing.

Rules and guards
- Tests/README.md "Test rules", mirrored in AGENTS.md, AGENT_START.md,
  CLAUDE.md, the agent contract's tests area, the /tests command, and a
  new test-writer agent that writes tests from the promise, not the code.
- scripts/dev/check-test-shape.py: ratchet that blocks NEW tests reading
  Sources/ as text or asserting on wall-clock time. Existing ones are
  grandfathered in .agents/test-shape-baseline.json, which only shrinks.
- Tests/quarantine.txt: bench a flaky fast-test suite the same day;
  run-tests.sh skips and lists it, the guard validates the file.
- scripts/dev/explain-missing-sources.py: when a hand-kept source list is
  missing a file, run-tests.sh and the Parakeet lifecycle smoke now print
  the file that declares the missing name and the list to add it to
  (replays the Sept 27 integration-smoke failure correctly).
- scripts/dev/check-known-traps.py: every Tools package runs in CI and has
  a matrix rule; every root command is documented.
- run-tests.sh defaults TZ to America/Chicago when unset.

One command
- check.sh: the mapped checks for your diff (agent-check.py), or the
  quick / full / hardware tiers, with plain PASS/FAIL and rerun commands.

Output and real-world layers
- MeetingMarkdownGoldenTests + Tests/Fixtures/golden-meetings: whole-file
  approved copies of saved meeting Markdown, time-zone proof.
- scripts/ops/nightly-hardware-smokes.sh: runs check.sh hardware plus the
  synthetic audio pass on a schedule (not installed; see the doc).
- scripts/dev/concurrency-census.sh: typecheck-only strict-concurrency
  census; 121 warnings across 8 Sources folders, baseline only shrinks.

Tests converted from wall-clock limits to outcome checks
- TranscriptedConstants detached timeout, ModelLoadProgressWaiter,
  DefaultInputDeviceMonitor submit, llama child shutdown, and the
  recovery-owner deadline (counts joins instead of timing the call).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds docs/testing-source-text-inventory.md: all 51 root fast-test files
that read Sources/*.swift as text, the promise each one guards, and a
verdict (converted, deleted, needs-seam with the smallest seam, or keep).

Tests only, no Sources changes. Across 8 files, 12 source-text pins are
gone and 19 behavior assertions are in:

- AgentConnectionGuide: relocate the capture library between two reads
  of folderPathsText instead of grepping for `static var`.
- AuditRegressionCoverageContract: run the real paster with a
  non-frontmost target and with a failing dispatcher.
- ObservabilityLogWriter: prepare a pre-existing 0644 log through
  ObservabilityLogFilePreparation and check it comes back 0600 intact.
- FocusOrderContract: check settings page identifiers and the ⌘1-⌘5
  order through TranscriptedSettingsPage.
- StatusItemPresentation: render each MenuBarGlyph and check it is a
  template drawn in neutral ink.
- FailedMeetingPresentation: call retryDisabled for complete, missing,
  non-retryable, and silent audio.
- ClipboardRestoringTextPaster, OverlayScreenSharePrivacy: drop pins
  already covered by existing behavior suites.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ps linux-checks already runs

main's #1874 added two source-text uses to DictationInputDeviceSelectionPolicyTests
before the guard existed, so the baseline takes them as grandfathered. #1821 made
repo-hygiene run linux-checks.sh, which already includes the new checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…allowlist test

The concurrency census flagged Sources/UI going from 53 to 54 warnings after
main's Today screen landed. Fixed all three Today warnings instead of
grandfathering: the two Int limits on TodayViewModel are nonisolated
constants, and the preview's arrowButton helper is @mainactor. UI is now 51,
total 119, and the baseline is lowered to match.

SentryEventPolicy.allowedDiagnosticTagKeys is internal instead of private, so
"Every allowlisted Sentry tag key survives the sanitizer" iterates the real
set instead of parsing SentryEventPolicy.swift as text.

Also: --help for the nightly and census scripts stops at the header comment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflicts: AgentConnectionGuideTests and FocusOrderContractTests, where the
source-text checks had become behavior checks here and main extended the old
text checks for the Writing page. Kept the behavior checks and added main's
Writing expectations to them (the writing folder follows a relocated library;
six pages in ⌘1–⌘6 order). main's Writing PR added 13 source-text reads to
UIAutomationSurfaceContractTests before the guard existed; grandfathered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scripts/dev/mutation-probe.py breaks one Swift file under Sources/ on
purpose, one small change at a time (== / !=, < / <=, > / >=, && / ||,
true / false, return-bool, single-digit n -> n+1, one-line if negation),
runs a test command, and reports KILLED / SURVIVED / COMPILE-ERROR with a
mutation score and a JSON report under build/mutation/.

It is careful with the real file: refuses paths outside Sources/, untracked
or dirty files, and bytes that are not the git index copy; needs a green
baseline; restores the original after every mutant, on Ctrl-C/SIGTERM/
SIGHUP and at exit; keeps a backup under build/mutation/.backup; and takes a
per-checkout lock. Stdlib only. --list plans without running, --lines
re-checks survivors against a wider test command, and --self-test runs
offline (wired into linux-checks.sh SELF_TEST_SCRIPTS and the test matrix).

First real runs (seed 1, --max 20, each file's own fast test):
DictationReadinessWaitPolicy 17 killed / 3 survived (85%),
DictationInputDeviceSelectionPolicy 11 / 9 (55%). 8 of those 9 survivors
also survive the whole fast suite, including a flipped guard in
DictationInputDeviceBindingPolicy.requireSelection. Details and a monthly
cadence are in docs/mutation-testing.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d past

- a real mic selection passes requireSelection; only device ID 0 is rejected
- a Bluetooth-transport headset with a plain name (WH-1000XM5) counts as Bluetooth
- with the lid closed, an external mic the user chose is still the one used
- ready-start hard recoveries stop exactly when the recovery budget is spent
- the readiness refresh gives up at exactly the timeout (exact binary values)

Each test was checked red against the one-operator mutation it guards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same bug as #1883, in the other VM test: linux-checks.sh points TMPDIR deep
inside the checkout, which put vnc.py's self-test socket past the 104-byte
macOS cap, so 'serve never opened its socket' failed every linux-checks run
in a worktree while the standalone run passed. Use a short /tmp directory
when the temp root is long.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@r3dbars
r3dbars marked this pull request as ready for review September 28, 2026 16:58
@r3dbars
r3dbars merged commit 134b63a into main Sep 28, 2026
8 checks passed
@r3dbars
r3dbars deleted the claude/ai-unit-tests-theory-9ec0d0 branch September 28, 2026 16:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant