Skip to content

feat(queen): bees run in their own containers (runners), rebased with registry keys and runner-branch review - #525

Merged
dmitrii-f-t27 merged 2 commits into
fix/queen-worker-provider-and-prompt-sizefrom
feat/queen-bee-runners-v2
Oct 3, 2026
Merged

dmitrii-f-t27 merged 2 commits into
fix/queen-worker-provider-and-prompt-sizefrom
feat/queen-bee-runners-v2

Conversation

@dmitrii-f-t27

Copy link
Copy Markdown
Collaborator

Owner request (2026-10-02): run bees in separate containers so more of them can work. The swarm hit its ceiling, 29/29 active, which is the most one 24 GB service can hold (24 GB is the maximum for the service).

This is #506 (design and reasoning in that PR) rebased onto the deploy branch, plus three things the deploy branch needed since 2026-09-21:

  1. Runners take keys from the contributor registry. Since feat(queen): manage contributor keys and preserve account XP #522/feat(queen): an owner switches the model of all their keys of a provider at once #524 the Queen allocates from contributorWorkerCandidates: managed keys added in the account (negative index), disabled keys removed, and the owner's chosen model. workerProviderForKeyIndex(keyIndex, runtime) now resolves an order through those same candidates. Without it, a runner would refuse nvidia #25 and ignore the owner's model.
  2. The Queen judges a runner's branch. The review reads queen-N in the Queen's checkout: the diff, the head, the witness git show of the specs, and a criteria worktree cut at the commit. A runner leaves that branch only in queen_bundle, so every runner bee would have stayed in wait forever, unjudged. The sweep now imports a claimed row's bundle first (importRunnerBranch). If the runner cut from a newer base, it fetches origin; if a stale worktree from an earlier local attempt holds the branch name, it replaces that worktree. If the import fails, the row stays in wait and nothing is counted against it.
  3. The entrypoint in runner mode. With TRIOS_QUEEN_BEES_RUN_ELSEWHERE=on, the Queen's own memory does not limit the swarm. The width becomes TRIOS_QUEEN_MAX_WORKERS_CEILING, or the lane count if that is unset.

Rollout

  • The Queen gets TRIOS_QUEEN_BEES_RUN_ELSEWHERE=on and also runs as a runner herself (TRIOS_BEE_RUNNER_SECONDS=15, TRIOS_BEE_RUNNER_SLOTS=24), so her own container keeps working and there is never a window with no runner.
  • Plus runner services trios-bee-runner-1/-2 (same repo/image, own volume, no TRIOS_QUEEN_TICK_SECONDS) and TRIOS_QUEEN_MAX_WORKERS_CEILING = 35 keys × 2 lanes = 70.

Verification (Bun 1.3.6)

  • tsc --noEmit clean; biome shows only warnings that already exist on the base.
  • tests/api/queen-*.test.ts: 722 pass, 0 fail. That includes queen-runner.test.ts 9/9, with 4 new tests: registry-key resolution (managed, owner model, disabled → refuse) and branch import on real git repositories (newer base, already imported, no bundle, stale worktree holding the branch).
  • tests/pglive/queen-runner-live.test.ts + queen-contributor-keys-live.test.ts: 21/21 against PostgreSQL 17, including the claim race with eight runners on one row.
  • Entrypoint derivation exercised in both modes: local mode is min(lanes, memory); runner mode is the ceiling.

Supersedes #506.

🤖 Generated with Claude Code

dmitrii-f-t27 and others added 2 commits October 2, 2026 20:56
…onger one machine wide

A bee was a thread of the Queen. `dispatchBee` cut a worktree on her volume
and sent the turn to her own `/chat`, so every bee shared her memory and her
disk, and the widest the swarm could ever be was one container. Measured on the
deployed one, 2026-09-18: about a gigabyte a bee against a 24 GB limit, and a
50 GB volume holding every worktree. Credentials stopped being the limit the day
the key list was read without a counter; the container was the limit after that.

So decision and execution are separated. The Queen still chooses everything she
chose before - the issue, its boundary, the credential - and writes it down; a
runner, in its own container, takes the order and does the work. Add replicas
and the swarm is wider.

THE ROW IS THE PROTOCOL. `queen_dispatch` already says which issue is in
flight, under which boundary, on which credential, and a review sweep and two
reapers read it. A second queue beside it would be a second answer to one
question. With TRIOS_QUEEN_BEES_RUN_ELSEWHERE=on the Queen writes the row with
`queued_at` and the brief and runs nothing; the row is in flight from that
moment, so the boundary is held and the key is taken exactly as before. A runner
(TRIOS_BEE_RUNNER_SECONDS set) claims one with an UPDATE over a row picked FOR
UPDATE SKIP LOCKED, so two runners cannot take one order - proven against a
real PostgreSQL 16 with eight runners reaching for one row at once.

NO SECRET IN THE TABLE. The row carries `key_index` only. The runner resolves it
with `workerProviderForKeyIndex` against the same worker variables the Queen
reads; a runner whose variables differ refuses rather than reaching for the
next key along, which is a different account.

THE WORK STILL LEAVES WITHOUT A PUSH CREDENTIAL. A runner has no volume anyone
can fetch from and may be gone minutes later, so it writes the bundle of
base..queen-N into `queen_bundle`, and the export route serves it from there
when the branch is not in its own checkout. `bundleOfBranch` is the existing
bundle code, extracted so both can use it. (The export route file was already
out of biome's format on the base; the pre-commit hook reformatted it, which is
most of that file's diff.)

THE REAPERS HAD TO LEARN WHERE A BEE RUNS, or the first restart would have been
a disaster:
- The boot reaper buried every unfinished row, because a restart of THIS
  container killed every bee in it. Runner bees do not die with the Queen. It
  now takes only rows the Queen ran herself (`queued_at IS NULL`).
- A runner vouches for its bee by renewing `claimed_at` every fifteen seconds
  while it waits for the ending. The stall reaper releases a claimed row whose
  runner has been silent for ten minutes - a vanished container - and never
  salvages it, because nothing of that bee is on this disk.
- The runner resets `dispatched_at` when the turn really starts, so the
  two-hour rule measures a turn and not the time the order waited.

The shared half of starting a turn - cut the worktree, start the turn, hand
back `begin` so the stream is read only after the caller's row exists - is one
function, `cutAndStart`, used by both paths; the ledger around it differs (an
insert for the Queen, an update for a runner) and stays with each caller.
`markBeeRunningHere` from #505 moved into it with the rest of the start.

Off by default on both sides. A deployment may be a Queen, a runner, or both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… Queen judges a runner's branch

Rebased #506 onto the deploy branch, where keys and models also come from
the contributor registry since #522/#524. A runner resolved an order's
key_index from its environment only, so a key added in the account (negative
index) or an owner's model choice would have been refused or ignored; it now
resolves through the same candidates the Queen allocated from.

The review reads queen-N in the Queen's checkout - diff, head, the specs at
that commit, criteria in a worktree cut from it - and a runner left that
branch only in queen_bundle, so every runner bee would have waited forever
unjudged. The sweep now brings a claimed row's bundle into the checkout first,
fetching origin when the runner cut from a newer base, and replacing a stale
worktree of an earlier local attempt that holds the branch name.

With bees running elsewhere the Queen's own memory says nothing about the
swarm's width, so the entrypoint takes the operator ceiling (else the lane
count) instead of dividing this container's memory.
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

✅ Tests passed — 2483/2543

Suite Passed Failed Skipped
✅ agent 87/87 0 0
✅ build 9/9 0 0
✅ cdp-protocol 5/5 0 0
✅ eval 93/93 0 0
✅ server-agent 272/272 0 0
✅ server-api 1340/1399 0 59
✅ server-browser 6/6 0 0
✅ server-integration 10/11 0 1
✅ server-lib 279/279 0 0
✅ server-pglive 25/25 0 0
✅ server-root 68/68 0 0
✅ server-skills 31/31 0 0
✅ server-tools 244/244 0 0
✅ shared 14/14 0 0

View workflow run

@dmitrii-f-t27
dmitrii-f-t27 merged commit 51138f0 into fix/queen-worker-provider-and-prompt-size Oct 3, 2026
17 of 18 checks passed
dmitrii-f-t27 pushed a commit that referenced this pull request Oct 3, 2026
The production branch now has its own runners (#525): operator replicas
that claim rows the Queen queued (queued_at / claimed_by) straight from
the database, with the operator's keys. Volunteer runners stay separate:
their offers never set queued_at, and operator rows never sit on a
volunteer lane.

Conflicts:
- queen-dispatch.ts reapers: a row is the container's to reap only when
  neither kind of runner holds it (queued_at IS NULL AND NOT_A_RUNNER).
  The stall reaper keeps the production branch's silent-runner clause
  for claimed rows and still skips volunteer lanes, which their lease
  reaper releases.
- queen-tick.ts imports, pg-migrate.ts columns: both kept.

dispatchBee still asks a volunteer runner first, then takes the
production path (a local bee, or a queued row for an operator replica).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
@github-actions
github-actions Bot deleted the feat/queen-bee-runners-v2 branch October 4, 2026 05:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant