Skip to content

feat(queen): a bee can run in its own container, so the swarm is no longer one machine wide - #506

Closed
dmitrii-f-t27 wants to merge 1 commit into
gHashTag:feat/queen-supervisorfrom
dmitrii-f-t27:feat/queen-bee-runners
Closed

dmitrii-f-t27 wants to merge 1 commit into
gHashTag:feat/queen-supervisorfrom
dmitrii-f-t27:feat/queen-bee-runners

Conversation

@dmitrii-f-t27

Copy link
Copy Markdown
Collaborator

Why

A bee is a thread of the Queen: dispatchBee cuts a worktree on her volume and sends the turn to her own /chat. Every bee shares her memory and her disk, so the widest the swarm can ever be is one container. Measured on the deployed one, 2026-09-18: about a gigabyte a bee against a 24 GB limit, and one 50 GB volume holding every worktree. Since #486 the key list has no upper edge, so the container is the limit that is left.

What changes

Decision and execution are separated. The Queen still chooses the issue, its boundary and the credential, and writes that down; a runner, in its own container, takes the order and does the work. Add replicas and the swarm is wider.

  • The row is the protocol. queen_dispatch already says which issue is in flight, under which boundary, on which credential, and the review sweep and both reapers read it. With TRIOS_QUEEN_BEES_RUN_ELSEWHERE=on the Queen writes the row with queued_at and the brief and runs nothing. The row is in flight from that moment, so the boundary is held and the key is taken exactly as before, and every reader keeps working without knowing where the bee runs.
  • One order, one runner. A runner claims with an UPDATE over a row picked FOR UPDATE SKIP LOCKED. Checked against a real PostgreSQL 16: eight runners reaching for one row at once, exactly one gets it.
  • No secret in the table. The row carries key_index; the runner resolves it with workerProviderForKeyIndex against the same worker variables the Queen reads. A runner whose variables differ refuses rather than taking the next key along, which is a different account.
  • The work still leaves without a push credential. A runner has no volume anyone can fetch from, so after the turn it writes the bundle of base..queen-N into a new queen_bundle table, and /queen/export/:issue serves it from there when the branch is not in its own checkout. The bundle code is the existing one, extracted as bundleOfBranch.
  • The shared half of a start is one function. cutAndStart cuts the worktree, starts the turn and hands back begin, so the stream is read only after the caller's row exists. The ledger around it differs, an insert for the Queen and an update for a runner, and stays with each caller. markBeeRunningHere from fix(queen): name where the server spins, and end a stalled loop in 3 min, not 12 #505 moved into it with the rest of the start.

The reapers had to learn where a bee runs

Without this the first Queen restart would have been a disaster.

  • Boot reaper. It released every unfinished row, because a restart of this container killed every bee in it. Runner bees do not die with the Queen. It now takes only rows the Queen ran herself, queued_at IS NULL.
  • A runner vouches for its bee. While it waits for the ending it renews claimed_at every 15 seconds. The stall reaper releases a claimed row whose runner has been silent for 10 minutes, a vanished container, and never salvages it, because nothing of that bee is on this disk.
  • The clock starts when the turn does. The runner resets dispatched_at when it actually starts the bee, so the two-hour rule measures a turn and not the time the order waited in the queue.

How to turn it on

Off by default on both sides; a deployment may be a Queen, a runner, or both.

  1. A second Railway service from this same repo and image, for example trios-bee-runner, with the replica count you want. No volume: each replica clones into its own disk at boot, as the entrypoint already does.
  2. Give it the same database and the same worker variables as the Queen (DATABASE_URL, TRIOS_QUEEN_WORKER_*, TRIOS_API_TOKEN, TRIOS_REPO_URL), plus TRIOS_BEE_RUNNER_SECONDS=15 and optionally TRIOS_BEE_RUNNER_SLOTS for bees per replica, default 1. Do not set TRIOS_QUEEN_TICK_SECONDS there.
  3. On the Queen: TRIOS_QUEEN_BEES_RUN_ELSEWHERE=on.

The container guard from #487 still applies, now inside each runner, which is where the memory actually is.

Verification

  • queen-dispatch.test.ts 101 pass, queen-runner.test.ts 5 pass (new), queen-resources 25, queen-volume-gc 13, queen-report-lines 9. New cases include a real runner start on a scratch repository with /chat stubbed, asserting the order is updated and not re-inserted and that the clock is reset.
  • The 39 api test files that import dispatch, tick, export or the runner: 587 pass, 61 skip, 4 fail. The same 4 fail on the base commit 269c4dc in my sparse checkout; nothing new.
  • Every new behavioural line was broken on purpose and a test caught it: salvage skip for runner rows, the clock reset, the queued-mode switch, the key-index inverse, the heartbeat's own-claim filter.
  • tests/pglive/queen-runner-live.test.ts (new, pglive group): the four claim cases passed against PostgreSQL 16 in Docker before I rebased onto fix(queen): name where the server spins, and end a stalled loop in 3 min, not 12 #505; the claim SQL did not change in the rebase. The two reaper cases in the same file I could not run here, because Docker Desktop on this machine stopped starting. They create the schema they need, so they should run on CI.
  • bun run typecheck clean; biome check shows only the warnings the base already has.

One pre-existing problem I found on the way, not fixed here: pg-migrate-live.test.ts reads information_schema under public, but createQueenPool sets search_path to trios, and a scratch database has no trios schema. So the migration fails silently there and the catalog comes back empty. That is why Tests / server-pglive fails on CI. My live test creates the schema first; the neighbour needs the same one line.

The export route file was already out of biome's format on the base, and the pre-commit hook reformatted it, which is most of that file's diff. The real change there is the bundleOfBranch extraction and the queen_bundle fallback.

🤖 Generated with Claude Code

…onger one machine wide

A bee was a thread of the Queen. `dispatchBee` cut a worktree on her volume
and sent the turn to her own `/chat`, so every bee shared her memory and her
disk, and the widest the swarm could ever be was one container. Measured on the
deployed one, 2026-09-18: about a gigabyte a bee against a 24 GB limit, and a
50 GB volume holding every worktree. Credentials stopped being the limit the day
the key list was read without a counter; the container was the limit after that.

So decision and execution are separated. The Queen still chooses everything she
chose before - the issue, its boundary, the credential - and writes it down; a
runner, in its own container, takes the order and does the work. Add replicas
and the swarm is wider.

THE ROW IS THE PROTOCOL. `queen_dispatch` already says which issue is in
flight, under which boundary, on which credential, and a review sweep and two
reapers read it. A second queue beside it would be a second answer to one
question. With TRIOS_QUEEN_BEES_RUN_ELSEWHERE=on the Queen writes the row with
`queued_at` and the brief and runs nothing; the row is in flight from that
moment, so the boundary is held and the key is taken exactly as before. A runner
(TRIOS_BEE_RUNNER_SECONDS set) claims one with an UPDATE over a row picked FOR
UPDATE SKIP LOCKED, so two runners cannot take one order - proven against a
real PostgreSQL 16 with eight runners reaching for one row at once.

NO SECRET IN THE TABLE. The row carries `key_index` only. The runner resolves it
with `workerProviderForKeyIndex` against the same worker variables the Queen
reads; a runner whose variables differ refuses rather than reaching for the
next key along, which is a different account.

THE WORK STILL LEAVES WITHOUT A PUSH CREDENTIAL. A runner has no volume anyone
can fetch from and may be gone minutes later, so it writes the bundle of
base..queen-N into `queen_bundle`, and the export route serves it from there
when the branch is not in its own checkout. `bundleOfBranch` is the existing
bundle code, extracted so both can use it. (The export route file was already
out of biome's format on the base; the pre-commit hook reformatted it, which is
most of that file's diff.)

THE REAPERS HAD TO LEARN WHERE A BEE RUNS, or the first restart would have been
a disaster:
- The boot reaper buried every unfinished row, because a restart of THIS
  container killed every bee in it. Runner bees do not die with the Queen. It
  now takes only rows the Queen ran herself (`queued_at IS NULL`).
- A runner vouches for its bee by renewing `claimed_at` every fifteen seconds
  while it waits for the ending. The stall reaper releases a claimed row whose
  runner has been silent for ten minutes - a vanished container - and never
  salvages it, because nothing of that bee is on this disk.
- The runner resets `dispatched_at` when the turn really starts, so the
  two-hour rule measures a turn and not the time the order waited.

The shared half of starting a turn - cut the worktree, start the turn, hand
back `begin` so the stream is read only after the caller's row exists - is one
function, `cutAndStart`, used by both paths; the ledger around it differs (an
insert for the Queen, an update for a runner) and stays with each caller.
`markBeeRunningHere` from gHashTag#505 moved into it with the rest of the start.

Off by default on both sides. A deployment may be a Queen, a runner, or both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@dmitrii-f-t27

Copy link
Copy Markdown
Collaborator Author

Superseded by #525 (rebased onto the deploy branch with registry keys and runner-branch review), merged 2026-10-03.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant