Repository navigation
Deflake MeetingPromptDetector tests with an explicit settle signal - #1906
Conversation
Signal pushes (mic, camera, output, requestEvaluation, EventKit change, workspace activate) each spawn an evaluate() task that awaits an off-main running-apps read. The tests waited a fixed ~100 ms for it, so on a loaded Mac they read the result before evaluation finished and got nil / 0 prompts. The detector now counts evaluations in flight (spawned passes, the first poll pass, title reads that re-evaluate, and timed re-checks once they fire) and exposes waitUntilEvaluationsSettle(). The tests await that instead of sleeping. The one test that waits on a 1 s timed re-check waits for its condition, then settles. Also: the DefaultInputDeviceNotificationLookupDispatcher ordering test used a 250 ms lookup timeout around a 50 ms fake lookup. Under load it timed out, delivered a nil device, and failed the self-write flag. The timeout isn't what that test checks, so it gets a 30 s bound.
The first loaded run (load avg ~280) failed "an unrecognized site waits before prompting" at the "not right away" check: with a 1 s wait counted from the mic push, a first evaluation pass slower than 1 s was already allowed to prompt. The hold-back check now uses an hour-long wait; a separate suite keeps the 1 s wait and checks that the detector's own re-check prompts.
|
Loaded run 1 ( |
|
Independent review (a separate Claude agent, full diff vs
Loaded run 1 on the split code: 138/138 at load average ~384. More runs in progress. |
|
Loaded repro on the final branch code ( Heads-up: this merged 27 s after #1907, which touches the same files. I read main's combined result. #1907's |
|
Loaded |
Why
MeetingPromptDetectorTestsfailed on and off on a loaded Mac (2 to 52 failures per run at load averages 60–150, cleanorigin/mainincluded). CI passed it every time. The failing checks all read the detector's result before its async evaluation had finished.Here's why. Every signal push (
updateMicInputUsers, camera, output,requestEvaluation, EventKit change, workspace activate) spawns aTaskthat runsevaluate(). That awaits an off-main running-apps read, and the tests don't stub it. The tests then waited a fixed ~100 ms (20 × yield + 5 ms sleep). On a busy machine that's not enough time.What changed
MeetingPromptDetectornow counts evaluations in flight and exposeswaitUntilEvaluationsSettle(). Counted: spawned passes (all signal pushes now go through onescheduleEvaluation()), the first poll pass afterstart(), browser title reads that re-evaluate when they land, and a timed re-check once its sleep ends. Sleeping re-checks and the 120 s poll aren't counted. The EventKit observer already runs on.main, so it now schedules synchronously viaMainActor.assumeIsolated, which means a post can't slip past the counter. There's no change to runtime behavior.extraMilliseconds:waits keep their sleep because it's a minimum gap so a title re-read is allowed. Then they settle. The one test that waits on a real 1 s re-check waits for its condition and settles, with a ~5 s give-up. Deflake the Meet-after-Not-now prompt test #1901'suntil:workarounds are now plain settles. No wall-clock assertions (check-test-shape.pyis clean).DefaultInputDeviceMonitorTests, the ordering test: its 250 ms lookup timeout around a 50 ms fake lookup timed out under load. It then delivered a nil device, so both observers sawfalse. That timeout isn't what the test checks, so it gets a 30 s bound. The assertion is unchanged.Checks
bash run-tests.sh --filter MeetingPromptDetector: 138/138 on an idle machine.check-test-shape.pyandcheck-source-pins.py --changed-onlyboth pass.yes× 3 per core, load average ~230–300) is still building locally. I'll post the results here.bash check.sh, independent review.🤖 Generated with Claude Code