feat(queen): a bee can run in its own container, so the swarm is no longer one machine wide - #506
Closed
dmitrii-f-t27 wants to merge 1 commit into
Closed
dmitrii-f-t27 wants to merge 1 commit into
dmitrii-f-t27 wants to merge 1 commit into
Conversation
…onger one machine wide A bee was a thread of the Queen. `dispatchBee` cut a worktree on her volume and sent the turn to her own `/chat`, so every bee shared her memory and her disk, and the widest the swarm could ever be was one container. Measured on the deployed one, 2026-09-18: about a gigabyte a bee against a 24 GB limit, and a 50 GB volume holding every worktree. Credentials stopped being the limit the day the key list was read without a counter; the container was the limit after that. So decision and execution are separated. The Queen still chooses everything she chose before - the issue, its boundary, the credential - and writes it down; a runner, in its own container, takes the order and does the work. Add replicas and the swarm is wider. THE ROW IS THE PROTOCOL. `queen_dispatch` already says which issue is in flight, under which boundary, on which credential, and a review sweep and two reapers read it. A second queue beside it would be a second answer to one question. With TRIOS_QUEEN_BEES_RUN_ELSEWHERE=on the Queen writes the row with `queued_at` and the brief and runs nothing; the row is in flight from that moment, so the boundary is held and the key is taken exactly as before. A runner (TRIOS_BEE_RUNNER_SECONDS set) claims one with an UPDATE over a row picked FOR UPDATE SKIP LOCKED, so two runners cannot take one order - proven against a real PostgreSQL 16 with eight runners reaching for one row at once. NO SECRET IN THE TABLE. The row carries `key_index` only. The runner resolves it with `workerProviderForKeyIndex` against the same worker variables the Queen reads; a runner whose variables differ refuses rather than reaching for the next key along, which is a different account. THE WORK STILL LEAVES WITHOUT A PUSH CREDENTIAL. A runner has no volume anyone can fetch from and may be gone minutes later, so it writes the bundle of base..queen-N into `queen_bundle`, and the export route serves it from there when the branch is not in its own checkout. `bundleOfBranch` is the existing bundle code, extracted so both can use it. (The export route file was already out of biome's format on the base; the pre-commit hook reformatted it, which is most of that file's diff.) THE REAPERS HAD TO LEARN WHERE A BEE RUNS, or the first restart would have been a disaster: - The boot reaper buried every unfinished row, because a restart of THIS container killed every bee in it. Runner bees do not die with the Queen. It now takes only rows the Queen ran herself (`queued_at IS NULL`). - A runner vouches for its bee by renewing `claimed_at` every fifteen seconds while it waits for the ending. The stall reaper releases a claimed row whose runner has been silent for ten minutes - a vanished container - and never salvages it, because nothing of that bee is on this disk. - The runner resets `dispatched_at` when the turn really starts, so the two-hour rule measures a turn and not the time the order waited. The shared half of starting a turn - cut the worktree, start the turn, hand back `begin` so the stream is read only after the caller's row exists - is one function, `cutAndStart`, used by both paths; the ledger around it differs (an insert for the Queen, an update for a runner) and stays with each caller. `markBeeRunningHere` from gHashTag#505 moved into it with the rest of the start. Off by default on both sides. A deployment may be a Queen, a runner, or both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Collaborator
Author
|
Superseded by #525 (rebased onto the deploy branch with registry keys and runner-branch review), merged 2026-10-03. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
A bee is a thread of the Queen:
dispatchBeecuts a worktree on her volume and sends the turn to her own/chat. Every bee shares her memory and her disk, so the widest the swarm can ever be is one container. Measured on the deployed one, 2026-09-18: about a gigabyte a bee against a 24 GB limit, and one 50 GB volume holding every worktree. Since #486 the key list has no upper edge, so the container is the limit that is left.What changes
Decision and execution are separated. The Queen still chooses the issue, its boundary and the credential, and writes that down; a runner, in its own container, takes the order and does the work. Add replicas and the swarm is wider.
queen_dispatchalready says which issue is in flight, under which boundary, on which credential, and the review sweep and both reapers read it. WithTRIOS_QUEEN_BEES_RUN_ELSEWHERE=onthe Queen writes the row withqueued_atand the brief and runs nothing. The row is in flight from that moment, so the boundary is held and the key is taken exactly as before, and every reader keeps working without knowing where the bee runs.UPDATEover a row pickedFOR UPDATE SKIP LOCKED. Checked against a real PostgreSQL 16: eight runners reaching for one row at once, exactly one gets it.key_index; the runner resolves it withworkerProviderForKeyIndexagainst the same worker variables the Queen reads. A runner whose variables differ refuses rather than taking the next key along, which is a different account.base..queen-Ninto a newqueen_bundletable, and/queen/export/:issueserves it from there when the branch is not in its own checkout. The bundle code is the existing one, extracted asbundleOfBranch.cutAndStartcuts the worktree, starts the turn and hands backbegin, so the stream is read only after the caller's row exists. The ledger around it differs, an insert for the Queen and an update for a runner, and stays with each caller.markBeeRunningHerefrom fix(queen): name where the server spins, and end a stalled loop in 3 min, not 12 #505 moved into it with the rest of the start.The reapers had to learn where a bee runs
Without this the first Queen restart would have been a disaster.
queued_at IS NULL.claimed_atevery 15 seconds. The stall reaper releases a claimed row whose runner has been silent for 10 minutes, a vanished container, and never salvages it, because nothing of that bee is on this disk.dispatched_atwhen it actually starts the bee, so the two-hour rule measures a turn and not the time the order waited in the queue.How to turn it on
Off by default on both sides; a deployment may be a Queen, a runner, or both.
trios-bee-runner, with the replica count you want. No volume: each replica clones into its own disk at boot, as the entrypoint already does.DATABASE_URL,TRIOS_QUEEN_WORKER_*,TRIOS_API_TOKEN,TRIOS_REPO_URL), plusTRIOS_BEE_RUNNER_SECONDS=15and optionallyTRIOS_BEE_RUNNER_SLOTSfor bees per replica, default 1. Do not setTRIOS_QUEEN_TICK_SECONDSthere.TRIOS_QUEEN_BEES_RUN_ELSEWHERE=on.The container guard from #487 still applies, now inside each runner, which is where the memory actually is.
Verification
queen-dispatch.test.ts101 pass,queen-runner.test.ts5 pass (new),queen-resources25,queen-volume-gc13,queen-report-lines9. New cases include a real runner start on a scratch repository with/chatstubbed, asserting the order is updated and not re-inserted and that the clock is reset.tests/pglive/queen-runner-live.test.ts(new, pglive group): the four claim cases passed against PostgreSQL 16 in Docker before I rebased onto fix(queen): name where the server spins, and end a stalled loop in 3 min, not 12 #505; the claim SQL did not change in the rebase. The two reaper cases in the same file I could not run here, because Docker Desktop on this machine stopped starting. They create the schema they need, so they should run on CI.bun run typecheckclean;biome checkshows only the warnings the base already has.One pre-existing problem I found on the way, not fixed here:
pg-migrate-live.test.tsreadsinformation_schemaunderpublic, butcreateQueenPoolsetssearch_pathtotrios, and a scratch database has notriosschema. So the migration fails silently there and the catalog comes back empty. That is whyTests / server-pglivefails on CI. My live test creates the schema first; the neighbour needs the same one line.The export route file was already out of biome's format on the base, and the pre-commit hook reformatted it, which is most of that file's diff. The real change there is the
bundleOfBranchextraction and thequeen_bundlefallback.🤖 Generated with Claude Code