Skip to content

fix(queen): ask the container before starting a bee, and give a bee's session back when its turn ends - #487

Merged
gHashTag merged 2 commits into
gHashTag:feat/queen-supervisorfrom
dmitrii-f-t27:feat/queen-resource-guard
Sep 19, 2026
Merged

gHashTag merged 2 commits into
gHashTag:feat/queen-supervisorfrom
dmitrii-f-t27:feat/queen-resource-guard

Conversation

@dmitrii-f-t27

Copy link
Copy Markdown
Collaborator

Why

Nothing in dispatch asks the container whether another bee fits. The worker ceiling is a number, the credential list is a number, and the one thing that actually decides how many bees run at once was never consulted.

Measured on 2026-09-17 from the Railway metrics of trios-agent-server: the container is limited to 24 GB and held 14.7 GB with sixteen bees running, about 0.85 GB a bee on top of roughly 1 GB idle. The volume is 50 GB with 29 to 31 in use and a peak of 35.8 that day. The same day the ceiling stopped being a constant (#485, #486) and the credential list reached 36 lanes. A full swarm would ask for about 31 GB of a 24 GB container, and the kernel answers that by killing the container with every running bee in it, each holding the only copy of its unfinished turn.

As a stopgap my key site now holds TRIOS_QUEEN_MAX_WORKERS at what the metrics say the container carries. That is a guess made from outside, every change of it costs a restart, and it cannot see a burst. The decision belongs where the memory is.

What changes

  • queen-resources.ts (new): whether one more bee fits. Before every start it reads what the container holds that the kernel cannot give back (memory.current minus the file cache, both lists, minus reclaimable slab) against memory.max, and the free space on the workspace volume (df -Pk, the same subprocess and for the same reason as volumeUsedPercent). cgroup v2 first, then v1, then /proc/meminfo. Nothing is cached: every reading and every variable is taken on the call, so the line moves with the container and no restart is needed.
  • The rule. Memory: refuse when used + (young + 1) x perBee > limit - min(limit x 10%, 8 GB). Disk: refuse when free < perBeeDisk + min(total x 10%, 10 GB). Memory is judged first because it is the failure that takes everything with it, and one cause is reported. Unknown is a refusal, not room; a machine that is not Linux has no container to protect and passes.
  • Young bees. The tick starts bees back to back and a bee takes minutes to allocate what it will hold, so thirty starts in one round would each read the same nearly empty container and each be admitted. Every start inside the warm-up window (10 min) reserves a bee's memory the reading cannot show yet. The reservation is keyed by the turn and ends in closeDispatch, because /chat answers 200 and streams a provider refusal afterwards. It reserves memory; it caps nothing by count.
  • Where it sits. In dispatchBee, after the key refusals and before prepareWorktree, so a refusal costs a measurement and nothing else. The tick already breaks on a refused dispatch and the next round fires when a bee finishes, which is exactly when memory comes back.
  • Disk gets one reap before it refuses, because it is the one resource this server can free itself, and the gate now sits in front of the only reaper there was. That reap stops at the guard's own line, not at the collector's 55%.
  • The reaper no longer takes the tree of a running bee. reapWorktrees takes a protect set (queen-<issue> of rows with started = true AND finished_at IS NULL), checked before the dirty check. "The newest six are likely to be running" was true with four bees; with twenty, the seventh-newest tree is a bee's too, and a bee that has read a lot and written nothing yet has a clean tree. prepareWorktree's own reap gets the same set, asked for only when a reap is about to happen. That one mattered more than the new gate: it fires at 80% used, earlier than the guard refuses, and on a scratch repository it removed two trees registered as running. The gate's reap also protects the branch being dispatched, because a re-dispatch could reap its own tree and then cut it afresh with -B, dropping the previous attempt's unpushed commit. An unreadable registry reaps nothing instead of protecting nobody.
  • The report says why. A round queend allowed and the container refused used to read "Started nothing. No reason given." under the headline "nothing to do". The headline is served to a browser by /queen/needs-you, whose contract is no path and no connection detail, so it gets a label from a closed vocabulary (no room in the container: memory or disk). The body gets the guard's summary, which carries every number, stays under 200 characters, and travels as a field rather than being cut out of prose that names a path.
  • It says what it reads, once. Queen resource guard is on {reads: "memory from cgroup v2, limit 24.0 GB from cgroup", volume: ...} when it first runs and again only when that changes, so the first thing to check after a deploy is in the log before anything is ever refused.

One choice you may want reversed

A container refusal is not written with recordDispatch, unlike the other refusals. It says nothing about the issue, and once the swarm is memory-bound it is how nearly every round ends: a row per round would add a queen_dispatch_history snapshot each time and rewrite the issue's live row, wiping the last attempt's conversation, tokens and outcome. The warning line and the round report carry it. One reviewer argued the opposite, that it should follow the existing convention. If you prefer the row, it is one call in dispatchBee.

Variables, all optional

Variable Default Meaning
TRIOS_QUEEN_RESOURCE_GUARD on off dispatches without the guard, and says so in the log
TRIOS_QUEEN_BEE_MEMORY_MB 1024 MiB one bee is expected to hold (measured 0.85 GB)
TRIOS_QUEEN_MEMORY_HEADROOM_PERCENT 10 kept free, capped at 8 GB
TRIOS_QUEEN_BEE_WARMUP_SECONDS 600 how long a started bee is reserved for
TRIOS_QUEEN_BEE_DISK_MB 2048 MiB one worktree is expected to take
TRIOS_QUEEN_DISK_HEADROOM_PERCENT 10 kept free, capped at 10 GB
TRIOS_QUEEN_MEMORY_LIMIT_MB unset states the limit where the cgroup says max. It can only lower the line: the kernel's own limit wins, and says so once in the log

A value that cannot be used is replaced by its default and said so, once. Numbers are printed in decimal GB, the unit Railway sells and the cgroup file uses.

Checked on Railway, not assumed

I could not open a shell in trios-agent-server, so I read the same files from inside another container on the same platform (my key site prints them at start): /sys/fs/cgroup/memory.max is 24000000000, memory.current and memory.stat are readable, there are no cgroup v1 files, and /proc/meminfo describes the host (MemTotal 338 GB). So on Railway the guard reads the real limit with no variable set. It also means the /proc/meminfo fallback measures the machine, which is right on a machine with no container around the server and wrong to hold against a stated container limit, so that combination answers "unknown" with a sentence saying why.

What I could not check from outside: how much active_file and reclaimable slab the Queen's container accumulates over a day. The guard does not count them either way. If you want the numbers: grep -E '^(anon|active_file|inactive_file|slab_reclaimable) ' /sys/fs/cgroup/memory.stat.

The second commit, measured after the first was written

The guard says "asked again every round, and a finishing bee is what gives memory back". On the deployed container that was false, and the metrics say why. Memory did not follow the bees; it followed the ENDINGS:

01:41  3 bees   1.8 GB      02:00  3 bees  10.5 GB
01:47  5 bees   2.0 GB      02:02  2 bees  11.1 GB
01:53  6 bees   2.5 GB      02:06  1 bee   11.0 GB
01:55  5 bees   5.0 GB      02:10  0 bees  11.0 GB

Seven bees cost 2.5 GB between them; an hour later, with no bee running, the container held eleven and gave none of it back. An earlier round the same night reached 21.9 GB of 24 and was saved by a restart.

Every dispatch gets a fresh conversation id, /chat builds an AgentSession for it and keeps it in the SessionStore, and nothing in queen-dispatch.ts ever asked for it back. The session holds the agent with the whole turn's message history - for a 65k-context coding turn, more than a gigabyte - so the baseline grew with every bee that finished, and only a redeploy cleared it. That is also why my site's watchdog kept cutting the ceiling (18 bees to 9) and it bought nothing: it was charging the bees for memory that was never theirs.

closeDispatch now calls DELETE /chat/:conversationId - the route that disposes the agent and drops it from the store, which has existed all along for the UI - over the same loopback the dispatch already uses to start a turn, before the refill signal. A turn that is over needs none of it: the review reads the transcript from queen_transcript, and a bee's conversation id is used once and never again. After the fix went in on my side by restart, the container idles at 0.5 GB instead of 11.

What this is not

  • Admission control only. A bee that was admitted and then grows far past its reserve can still take the container down; the headroom is the only buffer.
  • It caps nothing by count. The policy and the credentials still decide how many bees may run; this decides whether the next one fits. The Swift copies are untouched.
  • workerCapacityBreakdown keeps its closed three-integer shape (chore: bump version browseros-ai/BrowserOS#1308) and /queen/status is unchanged. A status that still says healthy_idle while the guard refuses is a fair follow-up, in its own change.

Seen on the way, not fixed here

Readers of the dispatch path found three things that make the memory baseline grow with uptime, each worth its own issue: finished bees' sessions are never deleted from the session store; a stall-reaped turn is not aborted, because startTurn holds no abort handle; the bash timeout kills su and not the process group.

Verification

  • The five test files of this change: 163 pass, 0 fail (queen-resources 25, queen-dispatch 91, queen-volume-gc 13, queen-report-lines 9, queen-round 25). The reaper and the re-dispatch cases run against a real scratch repository; two cases drive dispatchBee to a real start with /chat stubbed.
  • This time I built queend locally (swift build -c release --product queend), so the policy-gated round cases ran instead of skipping. The 32 test files that import dispatch, tick, report lines or public status: 428 pass, 13 skip, 7 fail with queend. The 7 are the same ones that fail on the base commit from my sparse checkout.
  • Every new behavioural test was checked by breaking the line it guards and watching it fail: the continue after keptRunning.push, the protect lookup, the running hand-off in prepareWorktree, noteBeeStarted, noteBeeEnded, the young count in the judgement, the second judgement after a reap, the own-branch protection, the reap's low mark, the report wiring in queen-tick.ts.
  • Three review rounds by independent agents. The first two found sixteen defects, all fixed and pinned above; the largest were the unprotected reap, the re-dispatch reaping its own tree, dead starts holding reservations, and active cache counted as used.
  • bun run typecheck clean; biome check shows only the three complexity warnings that exist on the base commit.

🤖 Generated with Claude Code

dmitrii-f-t27 and others added 2 commits September 17, 2026 21:53
…ee fits, so ask before every start

The worker ceiling is a number and the credential list is a number, and the one
thing that actually decides how many bees run at once - the container - was
never asked. Measured on 2026-09-17 from the Railway metrics: the container is
limited to 24 GB and held 14.7 GB with sixteen bees running, about 0.85 GB a
bee over 1 GB idle; the volume is 50 GB with 29 to 31 in use and 35.8 at the
day's peak. The same day the ceiling stopped being a constant and the credential
list reached 36 lanes. A full swarm would ask for about 31 GB of 24, and the
kernel answers that by killing the container with every running bee in it, each
holding the only copy of its unfinished turn.

queen-resources.ts is whether one more bee fits. Before every start it reads
what the container holds that the kernel cannot give back - memory.current
minus the file cache, both lists, minus reclaimable slab - against memory.max,
and the free space on the workspace volume. cgroup v2, then v1, then
/proc/meminfo. Nothing is cached, so the line moves with the container and with
the operator's variables and no restart is needed. Memory refuses when
used + (young + 1) x perBee > limit - min(10%, 8 GB); disk when
free < perBeeDisk + min(10%, 10 GB). Memory first, one cause reported. Unknown
is a refusal; a machine that is not Linux has no container to protect.

Young bees: the tick starts bees back to back and a bee takes minutes to
allocate what it will hold, so every start inside the warm-up reserves a bee's
memory the reading cannot show yet. The reservation is keyed by the turn and
ends in closeDispatch, because /chat answers 200 and streams a provider refusal
afterwards: kept as timestamps, a provider outage reserved ten GB for bees that
were already dead.

The gate sits in dispatchBee after the key refusals and before the worktree is
cut. Disk gets one reap before it refuses, bounded at the guard's own line
rather than the collector's 55%. A container refusal is not booked against the
issue: it says nothing about the issue and is how a memory-bound round ends.

The reaper no longer takes the tree of a running bee. reapWorktrees takes a
protect set and prepareWorktree's own reap gets it too - that one fires at 80%
used, earlier than the guard refuses, and on a scratch repository it removed two
trees registered as running. The gate's reap also protects the branch being
dispatched: a re-dispatch could reap its own tree and then cut it afresh with
-B, dropping the last attempt's unpushed commit. An unreadable registry reaps
nothing instead of protecting nobody.

The round report says why. /queen/needs-you serves the headline to a browser and
promises no path, so the headline gets a label from a closed vocabulary (no room
in the container: memory | disk) and the body gets the guard's summary, carried
as a field and never cut out of prose that names a path.

Checked from inside a Railway container rather than assumed: memory.max reads
24000000000, memory.current and memory.stat are readable, there is no cgroup v1,
and /proc/meminfo describes the host (338 GB). So a stated
TRIOS_QUEEN_MEMORY_LIMIT_MB can only lower the line - the kernel's limit wins -
and the host is never held against a stated container limit. Numbers are printed
in decimal GB, the unit the platform sells and the cgroup file uses.

Admission control only: a bee admitted and then grown far past its reserve can
still take the container down. It caps nothing by count; the Swift copies, the
closed capacity breakdown (browseros-ai#1308) and /queen/status are untouched.

Two review rounds found and these tests now pin: the unprotected reap, the
re-dispatch reaping its own tree, dead starts holding reservations, a stated
limit above the kernel's, active cache counted as used, a missing checkout read
as a full disk. Each new behavioural test was checked by breaking the line it
guards.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…d with agents nobody was running

MEASURED ON THE DEPLOYED CONTAINER, 2026-09-18, from the Railway metrics.
Memory did not follow the bees; it followed the ENDINGS:

  01:41  3 bees   1.8 GB      02:00  3 bees  10.5 GB
  01:47  5 bees   2.0 GB      02:02  2 bees  11.1 GB
  01:53  6 bees   2.5 GB      02:06  1 bee   11.0 GB
  01:55  5 bees   5.0 GB      02:10  0 bees  11.0 GB

Seven bees cost 2.5 GB between them. An hour later, with no bee running at
all, the container held eleven and gave none of it back. An earlier round the
same night reached 21.9 GB of 24 and was saved by a restart. Nothing here was
about how many bees run at once, which is why the ceiling the operator's
watchdog kept cutting - eighteen bees to nine - bought nothing: it was charging
the bees for memory that was never theirs.

Every dispatch gets a fresh conversation id, `/chat` builds an `AgentSession`
for it and keeps it in the SessionStore, and nothing in this file ever asked
for it back. The session holds the agent with the whole turn's message history;
for a 65k-context coding turn that is more than a gigabyte. So the baseline
grew with every bee that FINISHED, and only a redeploy cleared it.

`DELETE /chat/:conversationId` disposes the agent and drops it from the store.
It has existed all along, for the UI. `closeDispatch` now calls it, over the
same loopback the dispatch already uses to start a turn, before the refill
signal - so the memory is back before the next bee is started against it.

A turn that is over needs none of it: the review reads the transcript from
`queen_transcript`, and a bee's conversation id is used once and never again.
Failing to release is logged and nothing more; the ending stands either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gHashTag
gHashTag merged commit 1770c3b into gHashTag:feat/queen-supervisor Sep 19, 2026
12 of 18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants