fix(queen): name where the server spins, and end a stalled loop in 3 min, not 12 - #505
Merged
Merged
Conversation
…min, not 12 Measured 2026-09-21 14:44: the agent server spun one core at 100% for twelve minutes (cpu 973 s -> 1512 s), answered nothing, and the swarm stood until the liveness loop's twelfth missed /health probe ended it. Nothing recorded where it spun: pino writes asynchronously, so the lines logged just before a spin sit in a buffer the blocked thread never flushes. lib/stall-watch.ts: every log line, every tool call and each compaction step is written into a SharedArrayBuffer ring (no I/O). A Worker thread watches the main thread's heartbeat; when it stops for 15 s the worker writes the last twelve activities straight to fd 2 ([stall] lines), and says how long the stall lasted when the loop turns again. docker-entrypoint.sh: the main thread touches /tmp/trios-loop-heartbeat from a timer, which only fires when the event loop turns. A file older than LIVENESS_STALL (180 s) ends the server at once. Busy is still not dead: a loop slowed by twenty bees keeps touching the file. The twelve-probe /health rule stays as the backstop, and is alone in charge when the file is missing. Verified locally with bun 1.3.11 under the extracted run_supervised: a 25 s sync spin is reported with its last activities; a spinning server is ended on the stale heartbeat (exit 143) after one missed probe; a server kept 80% busy for 30 s is not touched. Compaction tests: 125 pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
✅ Tests passed — 2379/2439
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The swarm stood for twelve minutes today (2026-09-21 14:44-14:56 UTC): the agent server spun one core at 100% (cpu 973 s -> 1512 s over the liveness readings), answered nothing, and was ended only by the twelfth missed
/healthprobe. Nothing said where it spun - pino's asynchronous writes leave the last lines before a spin in an unflushed buffer.What changes
lib/stall-watch.ts+ worker. Log lines, tool calls (executeTool) and each compaction step are recorded in a SharedArrayBuffer ring, no I/O. A Worker thread watches the main thread's heartbeat and, after 15 s of silence, writes the last 12 activities to fd 2 as[stall]lines; when the loop turns again it logs how long the stall was (so short stalls under load become visible too).docker-entrypoint.sh. The main thread touches/tmp/trios-loop-heartbeatfrom a timer, which fires only when the event loop turns. A heartbeat older thanLIVENESS_STALL(180 s) ends the server at once. A server busy with twenty bees still turns its loop, so the reason the probe rule needed 12 minutes (it killed busy servers at 4 x 10 s) doesn't apply to this check. The 12-probe rule stays as the backstop and is the only rule when the file is missing.Verified
run_supervisedextracted from the entrypoint, with a server that spins after 6 s: ended on the stale heartbeat after one missed probe (exit 143), with the[stall]report in the same log.🤖 Generated with Claude Code