Skip to content

feat(queen): a second provider's keys had nowhere to go, because one URL was the whole pool - #484

Merged
gHashTag merged 3 commits into
gHashTag:feat/queen-supervisorfrom
dmitrii-f-t27:feat/queen-worker-endpoint-pools
Sep 17, 2026
Merged

gHashTag merged 3 commits into
gHashTag:feat/queen-supervisorfrom
dmitrii-f-t27:feat/queen-worker-endpoint-pools

Conversation

@dmitrii-f-t27

@dmitrii-f-t27 dmitrii-f-t27 commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

What was measured

Production, 2026-09-17: ten working Z.ai keys carried sixteen bees (/queen/status: capacity-full, 16 of 16) while three working NVIDIA keys sat unused beside them. The deployment has one TRIOS_QUEEN_WORKER_BASE_URL and one model, and every generic key is sent there. An NVIDIA key in TRIOS_QUEEN_WORKER_API_KEY_n would not widen the swarm: the bee holding it would send it to api.z.ai and get a 401 that arrives blamed on the work.

What changes

Numbered endpoint pools. The first pool is the variables that already exist, unchanged in name and meaning. Further pools start at 2:

Variable Meaning
TRIOS_QUEEN_WORKER_POOL_<n>_BASE_URL the endpoint (required)
TRIOS_QUEEN_WORKER_POOL_<n>_MODEL the model that endpoint serves (required, never inherited)
TRIOS_QUEEN_WORKER_POOL_<n>_PROVIDER openai-compatible (default) or zai
TRIOS_QUEEN_WORKER_POOL_<n>_CONTEXT that model's context window
TRIOS_QUEEN_WORKER_POOL_<n>_API_KEY, _API_KEY_2 … _API_KEY_1024 the pool's keys

Provider, base URL, model and context window travel to /chat from the same pool as the key. One allocator runs over the concatenated list, so a wave of bees takes one credential of every pool before any key carries a second lane.

queen_dispatch.key_index stays one integer. First-pool keys keep the index they have always had, so rows in flight across the deploy still mean the same key. A key of pool n is (n - 1) * POOL_KEY_STRIDE + position (pool 2 starts at 10000), using the pool number rather than its rank among today's valid pools, so disconnecting pool 2 does not rename every busy key of pool 3.

Second commit: the seventeenth key was read by nothing

keysFor() read the unsuffixed variable and then _2 … _16. Sixteen was never a measurement: it is where the loop happened to stop, equal to the policy ceiling on bees only by coincidence. The ceiling bounds how many bees run; the key list bounds how many credentials the rotation may spread them over, and a free key refuses its third concurrent request in under half a second (1302), so a swarm at the ceiling wants more keys than bees. A TRIOS_QUEEN_WORKER_API_KEY_17 was read by nothing and reported by nothing.

The suffix now runs to MAX_KEYS_PER_POOL = 1024 for every pool and for the legacy provider variables (they share keysFor). Widening the list does not widen the swarm: queenWorkerLimit() still answers at most 16, and a test pins seventeen credentials to a capacity of sixteen. POOL_KEY_STRIDE moves to 10000 with it, and a test holds MAX_KEYS_PER_POOL below the stride so the last key of a pool can never be the first key of the next. 1024 rather than a round thousand (third commit) so the key bound can never sit below a worker ceiling expressed as a power of two: with one lane per credential N bees need N keys. The worker ceiling itself is not changed in this pull request.

What deliberately does not change

  • A single endpoint takes the old code path byte for byte. All 66 existing dispatch tests pass untouched; 15 new ones cover pools and the key bound (81 total).
  • The closed breakdown (chore: bump version browseros-ai/BrowserOS#1308) stays three integers. Lanes per credential stay one number for the deployment, so it remains an exact statement: credentials × lanes, bounded by the policy ceiling. No pool name, URL or variable name enters it.
  • An explicit endpoint stays authoritative, including its refusal: a valid second pool does not take over from a first pool with no key.
  • A local Ollama stays one measured inference slot and reads no pools.
  • One secret named by two pools is one credential (fix(agent): install MCP against the proxy URL + bump agent-mcp-manager to v0.0.2 browseros-ai/BrowserOS#1293).
  • TRIOS_QUEEN_MAX_WORKERS and its ceiling of 16 are untouched; pools widen the credential list, not the policy.

A pool someone began to configure and dispatch ignored reads as configured in a variable editor and supplies nothing. endpointPoolProblems() names the missing variable (names only, never a value) and the "every key is busy" refusal carries it.

The dispatch detail printed key ${keyIndex + 1}; with a durable index of 1002 that would read "key 1003/3". It now prints the pool and the position inside it.

Verification

  • bun test tests/api/queen-dispatch.test.ts: 81 pass, 0 fail (66 before).
  • All 30 test files that import the dispatch, tick or public status code: 342 pass after, 330 before, and the same 11 failures in both runs. Those 11 come from my sparse checkout lacking t27-core, the Swift sources and the canonical tree; none touches this change.
  • bun run typecheck in apps/server: clean. biome check on both files: only the two pre-existing complexity warnings.
  • The patch applies cleanly to fix/queen-worker-provider-and-prompt-size (the branch production deploys from, 0e37000d) as well as to feat/queen-supervisor. That branch already carries the overload retry measured against integrate.api.nvidia.com, which a NIM pool will want.

To use it in production after merge

TRIOS_QUEEN_WORKER_POOL_2_BASE_URL=https://integrate.api.nvidia.com/v1
TRIOS_QUEEN_WORKER_POOL_2_MODEL=nvidia/nemotron-3-super-120b-a12b
TRIOS_QUEEN_WORKER_POOL_2_API_KEY=…   (three working keys exist, from three accounts)

Not measured here: how many concurrent requests one NIM key carries. The Z.ai figure of two lanes was measured on Z.ai only; NIM's free tier is about 40 requests a minute per account.

🤖 Generated with Claude Code

dmitrii-f-t27 and others added 3 commits September 17, 2026 17:16
…URL was the whole pool

Measured on 2026-09-17 in production: ten working Z.ai keys carried sixteen
bees while three working NVIDIA keys sat unused beside them. A key handed to a
bee is only as good as the URL the bee sends it to, and the deployment has
exactly one: TRIOS_QUEEN_WORKER_BASE_URL, one model, every generic key sent
there. Putting an NVIDIA key into TRIOS_QUEEN_WORKER_API_KEY_n would not widen
the swarm, it would hand a bee a 401 that arrives blamed on the work.

So there are now numbered endpoint pools. The FIRST pool is the variables that
already exist, unchanged in name and meaning. Further pools start at 2:

  TRIOS_QUEEN_WORKER_POOL_<n>_BASE_URL   the endpoint (required)
  TRIOS_QUEEN_WORKER_POOL_<n>_MODEL      the model that endpoint serves (required)
  TRIOS_QUEEN_WORKER_POOL_<n>_PROVIDER   openai-compatible (default) or zai
  TRIOS_QUEEN_WORKER_POOL_<n>_CONTEXT    that model's context window
  TRIOS_QUEEN_WORKER_POOL_<n>_API_KEY, _API_KEY_2 ... _API_KEY_16

What travels to /chat - provider, base URL, model, context window - comes from
the same pool as the key. One allocator runs over the concatenated list, so a
wave of bees takes one credential of every pool before any key carries a second
lane: the rule availableKeyIndex already applied inside one pool.

queen_dispatch.key_index is one integer, read back every tick to learn which
credentials are busy. First-pool keys keep the index they have always had, so
rows in flight across the deploy still mean the same key. A key of pool n is
(n - 1) * POOL_KEY_STRIDE + position (pool 2 starts at 1000). The pool NUMBER
is used rather than its rank among today's valid pools, so disconnecting pool 2
does not rename every busy key of pool 3.

Deliberately not changed:
- A single endpoint takes the old code path byte for byte; every existing test
  passes untouched (66 before, 78 after).
- The closed breakdown (browseros-ai#1308) stays three integers. Lanes per credential stay
  one number for the deployment, which keeps it an exact statement:
  credentials x lanes, bounded by the policy ceiling. No pool name, URL or
  variable name enters it.
- An explicit endpoint stays authoritative, including its refusal: a valid
  second pool does not take over from a first pool with no key.
- A local Ollama stays one measured inference slot and reads no pools.
- The model is required per pool and never inherited: the first pool's model
  name sent to another provider is a 404.
- One secret named by two pools is one credential (browseros-ai#1293).

A pool someone began to configure and dispatch ignored reads as configured in a
variable editor and supplies nothing. endpointPoolProblems() names the missing
variable (names only, never a value), and the "every key is busy" refusal
carries it, because that is the moment an operator asks why the swarm is
narrower than the variables suggest.

The dispatch detail printed `key ${keyIndex + 1}`; with a durable index of 1002
that would read "key 1003/3". It now prints the pool and the position in it.

Applies cleanly to fix/queen-worker-provider-and-prompt-size (the branch
production deploys from) as well as to feat/queen-supervisor.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…here a loop ended

keysFor() read the unsuffixed variable and then _2 ... _16. Sixteen was never a
measurement: it is where the loop happened to stop, and it equals the policy
ceiling on bees only by coincidence. They are different quantities. The ceiling
bounds how many bees RUN. The key list bounds how many credentials the rotation
may SPREAD them over - and a free key refuses its third concurrent request in
under half a second (1302), so a swarm at the ceiling wants more keys than
bees, not the same number.

A TRIOS_QUEEN_WORKER_API_KEY_17 was therefore read by nothing and reported by
nothing: it looks configured in a variable editor and supplies nothing, the
zero-length-key trap one step to the right.

The suffix now runs to MAX_KEYS_PER_POOL = 1000, for every pool and for the
legacy provider variables alike (they share keysFor). Widening the list does
not widen the swarm: queenWorkerLimit() still answers at most 16, and a test
pins seventeen credentials to a capacity of sixteen.

POOL_KEY_STRIDE moves from 1000 to 10000 with it. The durable index of pool n
starts at (n - 1) * stride, so a pool that could hold `stride` keys would have
its last key collide with the next pool's first. Nothing has been written with
the old stride - it exists only in this pull request - and a test now holds
MAX_KEYS_PER_POOL below POOL_KEY_STRIDE so the two cannot drift.

Dispatch tests: 81 pass (78 before this commit). Typecheck clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…g of 1024

MAX_KEYS_PER_POOL becomes 1024. With one lane per credential a swarm of N bees
needs N keys, so a key bound under the worker bound is a second, hidden ceiling
that reports nothing when it is hit. The operator is asking for the policy
ceiling to be raised to 1024; a round thousand would have been that hidden
ceiling the day it was. The worker ceiling itself is NOT changed here - it is a
policy with a Swift mirror, and raising it is a separate, deliberate decision.

Dispatch tests: 81 pass. Typecheck clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@gHashTag
gHashTag merged commit a05b474 into gHashTag:feat/queen-supervisor Sep 17, 2026
1 check failed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants