feat(queen): runner cabinet - lend a lane without lending a key - #518
dmitrii-f-t27 wants to merge 9 commits into
Conversation
A signed-in person (app.t27.ai session, verified by asking its issuer's whoami for telegram_id) can mint, list and revoke runner tokens at /queen/me/runners. A runner process on the lender's own machine presents that token at /queen/runner/heartbeat and learns its lane. The provider key never reaches the Queen: no route accepts, stores or returns one. - queen_runner table: only the token's SHA-256 and last four characters are kept; at most 5 live runners per person. - A runner's lane is 100000000 + id, a key_index block no operator pool reaches, so dispatch rows on it are credited by the existing leaderboard arithmetic. - Leaderboard gathers runner lanes by person (telegram_id), never by name, so a runner cannot merge into an operator's row, and never links a runner to an unverified GitHub login. - CORS for /queen/me/* is exactly https://app.t27.ai, bearer only, no credentials; both new mounts are on the route-guard allowlist with an own-bearer reason, and the audit pins are re-measured. Taking tasks and handing work back (claim/complete, remote review, runner CLI) is the next stage; the heartbeat says so instead of offering work. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
|
The same failure appears on #517 (two runs), which was merged. Nothing in this PR can fix it. It needs the token secret to be set, or the CLA workflow to be disabled for this fork. A re-run would fail the same way, so I have not re-run it. Generated by Claude Code |
✅ Tests passed — 2457/2517
|
|
All Every job hangs in Reproduced locally on
Two things on
This PR does not touch either file, so I have not widened it. Proposed patch (either part alone unblocks CI; both are better): git rm --cached trios/agent-server/apps/server/node_modules # drop the Mac-local symlink# .github/workflows/test.yml, "Setup Bun"
- name: Setup Bun
uses: oven-sh/setup-bun@v2
with:
bun-version-file: trios/agent-server/package.jsonOnce that lands on Generated by Claude Code |
…dules symlink Every `Tests / *` job on #517 and #518 was cancelled at the 20-minute budget while still inside `bun ci`, before any test ran. Two things combined: - trios/agent-server/apps/server/node_modules was committed as a symlink to a local macOS path. `.gitignore` said `node_modules/`, which only matches directories, so the symlink slipped through. - setup-bun ran without a version. The `packageManager: bun@1.3.6` pin lives in trios/agent-server/package.json, not at the repo root, so CI got the latest release (1.4.2), which hangs on that dangling symlink. 1.3.x installs past it. Untrack the symlink, make the ignore rule match files too, and read the Bun version from the workspace package.json in test.yml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…example.com With installs no longer hanging, server-tools ran for the first time since 2026-09-23 and failed one test: get_page_content read https://example.com 57 ms after opening it and found no "Example Domain". The test is about extracting text, so it now writes that text into about:blank with evaluate_script, as get_page_links already does. Locally (BrowserOS AppImage, headless, --no-sandbox): the old test fails the same way; the new one passes 3/3. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
… process
server-tools still exited 1 after every test in observation.test.ts
passed: before navigation-newtab-guard.test.ts the helper ran
`lsof -ti :<cdp port>`, which also lists clients still connected to the
port. One of them was the bun test process itself (its CDP socket to the
previous file's browser), so the SIGTERM ended the whole run and no
junit report was written ("workflow > server-tools setup").
Use `lsof -ti tcp:<port> -sTCP:LISTEN` and drop process.pid.
Locally, input.test.ts + navigation-newtab-guard.test.ts in one process:
before, exit 143 right after "Terminating process(es) <own pid>, ...";
after, 18 pass / 0 fail. The whole test:tools group now runs to the end
(242 pass; the 2 local failures load https://example.com, which this
sandbox's browser cannot reach and CI can).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…example.com With the run no longer killing itself, server-tools finished in CI with 243 pass / 1 fail: `wait_for finds text on page` waited its full 10 s for "Example Domain" on https://example.com and never saw it - the same page get_page_content could not read either. The page now adds that text itself 500 ms after load, so the test still proves wait_for waits, with nothing outside the runner involved. Locally: 2/2 wait_for tests pass on repeat; the whole test:tools group is 243 pass, the one local failure being take_screenshot (a 60 s hang in this sandbox only - it passes in CI). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
… too server-tools on 3d57649 ran clean except one test that had passed on both earlier runs: `search_dom > finds multiple elements with CSS class selector` (123 ms, fewer than 3 matches). It searches once, straight after new_page - the race this file already names and fixes with searchUntil for two sibling tests. Use the same helper here. Locally: search_dom 13/13, three runs in a row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…udget `the salvage commit > never splits a rename across the path cap` runs real git over 205 files and salvageWorktree. It takes ~2 s for the whole file locally and passed on the two CI runs before, then hit bun's 5 s default once on a loaded runner (job 110500921083) with nothing in the change touching salvage. A git-heavy fixture test should not share the budget of a pure unit test. Locally: queen-salvage-guards.test.ts 13 pass / 0 fail. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
Conflicts, both additive: - server.ts: keep the runner mounts and add /queen/contributor-keys. - queen-leaderboard.ts: rank() takes the operator map merged with contributorOwnerNames(), and the runner owners beside it. Also ports the route-guard fix from #520: #522 left /queen/contributor-keys out of the audit, which turns four route-guard tests red on the base. It is allowlisted with its own capability guard and the pins are re-measured (48 mounts; 25 /queen: 8/8/9). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…dules symlink (#520) * fix(ci): pin Bun to the workspace version and untrack a local node_modules symlink Every `Tests / *` job on #517 and #518 was cancelled at the 20-minute budget while still inside `bun ci`, before any test ran. Two things combined: - trios/agent-server/apps/server/node_modules was committed as a symlink to a local macOS path. `.gitignore` said `node_modules/`, which only matches directories, so the symlink slipped through. - setup-bun ran without a version. The `packageManager: bun@1.3.6` pin lives in trios/agent-server/package.json, not at the repo root, so CI got the latest release (1.4.2), which hangs on that dangling symlink. 1.3.x installs past it. Untrack the symlink, make the ignore rule match files too, and read the Bun version from the workspace package.json in test.yml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(tools): get_page_content reads a constructed page, not the live example.com With installs no longer hanging, server-tools ran for the first time since 2026-09-23 and failed one test: get_page_content read https://example.com 57 ms after opening it and found no "Example Domain". The test is about extracting text, so it now writes that text into about:blank with evaluate_script, as get_page_links already does. Locally (BrowserOS AppImage, headless, --no-sandbox): the old test fails the same way; the new one passes 3/3. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(helpers): killProcessOnPort kills listeners only, never the test process server-tools still exited 1 after every test in observation.test.ts passed: before navigation-newtab-guard.test.ts the helper ran `lsof -ti :<cdp port>`, which also lists clients still connected to the port. One of them was the bun test process itself (its CDP socket to the previous file's browser), so the SIGTERM ended the whole run and no junit report was written ("workflow > server-tools setup"). Use `lsof -ti tcp:<port> -sTCP:LISTEN` and drop process.pid. Locally, input.test.ts + navigation-newtab-guard.test.ts in one process: before, exit 143 right after "Terminating process(es) <own pid>, ..."; after, 18 pass / 0 fail. The whole test:tools group now runs to the end (242 pass; the 2 local failures load https://example.com, which this sandbox's browser cannot reach and CI can). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(tools): wait_for waits for text a data: page adds, not the live example.com With the run no longer killing itself, server-tools finished in CI with 243 pass / 1 fail: `wait_for finds text on page` waited its full 10 s for "Example Domain" on https://example.com and never saw it - the same page get_page_content could not read either. The page now adds that text itself 500 ms after load, so the test still proves wait_for waits, with nothing outside the runner involved. Locally: 2/2 wait_for tests pass on repeat; the whole test:tools group is 243 pass, the one local failure being take_screenshot (a 60 s hang in this sandbox only - it passes in CI). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(tools): the class-selector search_dom test retries the load race too server-tools on 3d57649 ran clean except one test that had passed on both earlier runs: `search_dom > finds multiple elements with CSS class selector` (123 ms, fewer than 3 matches). It searches once, straight after new_page - the race this file already names and fixes with searchUntil for two sibling tests. Use the same helper here. Locally: search_dom 13/13, three runs in a row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(queen): give the 205-file salvage rename test an explicit 30 s budget `the salvage commit > never splits a rename across the path cap` runs real git over 205 files and salvageWorktree. It takes ~2 s for the whole file locally and passed on the two CI runs before, then hit bun's 5 s default once on a loaded runner (job 110500921083) with nothing in the change touching salvage. A git-heavy fixture test should not share the budget of a pure unit test. Locally: queen-salvage-guards.test.ts 13 pass / 0 fail. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(queen): the route-guard audit knows /queen/contributor-keys #522 mounted /queen/contributor-keys and left the route-guard audit unchanged, so feat/queen-supervisor fails four route-guard tests: 46 mounts against a pin of 45, 23 /queen mounts against 22, and an unguarded mount nobody allowlisted. The route is a server-to-server door for the app render proxy and has its own guard: a bearer equal to QUEEN_CONTRIBUTOR_PROXY_TOKEN (32+ bytes, timingSafeEqual) plus a verified contributor header, and it is off while that token is unset. The trusted-origin check would refuse its only caller, so it is allowlisted with that reason and the pins are re-measured. No other number moved. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 --------- Co-authored-by: Claude <noreply@anthropic.com>
…516) * feat(queen): record what accepted .t27 work has earned, append-only The first half of spec authors mining TRI: nothing can be minted from a number nobody wrote down. Every Queen round now records each accepted turn whose boundary names a .t27 file as one earning per (repository, issue, judged commit), with work_id = sha256('t27-accept:v1|repo|issue|commit') so anyone can recompute it from public data. Why a table, when the leaderboard derives its score on read: a CI take-back edits queen_dispatch in place, so an acceptance derived on read would vanish instead of showing as taken back. Rows here are inserted and revoked, never deleted. A later sendBack/escalate of the same commit revokes an earning, and the revocation is final. GET /queen/public-earnings serves the record (public-read, no titles, no worker text, no notes) and says in its own body that nothing is withdrawable: no token is deployed, TRI per spec is undecided, and an accept does not yet require a merge. Tests: unit (grouping, query parameters, 503 without a database) and a live PostgreSQL test in tests/pglive covering idempotency, the work_id hash, the non-.t27 exclusion, take-back revocation and a new commit after a send-back. Route-guard census re-measured: 46 mounts, 9 public-read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ci): pin Bun to the workspace version and untrack a local node_modules symlink Every `Tests / *` job on #517 and #518 was cancelled at the 20-minute budget while still inside `bun ci`, before any test ran. Two things combined: - trios/agent-server/apps/server/node_modules was committed as a symlink to a local macOS path. `.gitignore` said `node_modules/`, which only matches directories, so the symlink slipped through. - setup-bun ran without a version. The `packageManager: bun@1.3.6` pin lives in trios/agent-server/package.json, not at the repo root, so CI got the latest release (1.4.2), which hangs on that dangling symlink. 1.3.x installs past it. Untrack the symlink, make the ignore rule match files too, and read the Bun version from the workspace package.json in test.yml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * feat(queen): one earning by work id, and the epoch-1 amount (27 TRI) GET /queen/public-earnings/:workId returns one earning with its earner's GitHub login, the scheme and TRI_PER_SPEC = 27 (owner decision O2, 2026-10-01). This is what each TRI signer reads before it signs; the merge rule (O4) is checked by the signers on GitHub, not here. Status text says what is true: mintable on TON testnet only, V1, signer quorum, NOT trustless. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(tools): get_page_content reads a constructed page, not the live example.com With installs no longer hanging, server-tools ran for the first time since 2026-09-23 and failed one test: get_page_content read https://example.com 57 ms after opening it and found no "Example Domain". The test is about extracting text, so it now writes that text into about:blank with evaluate_script, as get_page_links already does. Locally (BrowserOS AppImage, headless, --no-sandbox): the old test fails the same way; the new one passes 3/3. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * feat(queen): every earning credited to one GitHub login GET /queen/public-earnings/by/:github lists the earnings of the keys lent under that login, newest first: what the TRI wallet shows its owner as claimable. The ledger's 'recent' is capped at 100 and cannot serve this. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(helpers): killProcessOnPort kills listeners only, never the test process server-tools still exited 1 after every test in observation.test.ts passed: before navigation-newtab-guard.test.ts the helper ran `lsof -ti :<cdp port>`, which also lists clients still connected to the port. One of them was the bun test process itself (its CDP socket to the previous file's browser), so the SIGTERM ended the whole run and no junit report was written ("workflow > server-tools setup"). Use `lsof -ti tcp:<port> -sTCP:LISTEN` and drop process.pid. Locally, input.test.ts + navigation-newtab-guard.test.ts in one process: before, exit 143 right after "Terminating process(es) <own pid>, ..."; after, 18 pass / 0 fail. The whole test:tools group now runs to the end (242 pass; the 2 local failures load https://example.com, which this sandbox's browser cannot reach and CI can). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(tools): wait_for waits for text a data: page adds, not the live example.com With the run no longer killing itself, server-tools finished in CI with 243 pass / 1 fail: `wait_for finds text on page` waited its full 10 s for "Example Domain" on https://example.com and never saw it - the same page get_page_content could not read either. The page now adds that text itself 500 ms after load, so the test still proves wait_for waits, with nothing outside the runner involved. Locally: 2/2 wait_for tests pass on repeat; the whole test:tools group is 243 pass, the one local failure being take_screenshot (a 60 s hang in this sandbox only - it passes in CI). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(tools): the class-selector search_dom test retries the load race too server-tools on 3d57649 ran clean except one test that had passed on both earlier runs: `search_dom > finds multiple elements with CSS class selector` (123 ms, fewer than 3 matches). It searches once, straight after new_page - the race this file already names and fixes with searchUntil for two sibling tests. Use the same helper here. Locally: search_dom 13/13, three runs in a row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(queen): give the 205-file salvage rename test an explicit 30 s budget `the salvage commit > never splits a rename across the path cap` runs real git over 205 files and salvageWorktree. It takes ~2 s for the whole file locally and passed on the two CI runs before, then hit bun's 5 s default once on a loaded runner (job 110500921083) with nothing in the change touching salvage. A git-heavy fixture test should not share the budget of a pure unit test. Locally: queen-salvage-guards.test.ts 13 pass / 0 fail. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * test(queen): the route-guard audit knows /queen/contributor-keys #522 mounted /queen/contributor-keys and left the route-guard audit unchanged, so feat/queen-supervisor fails four route-guard tests: 46 mounts against a pin of 45, 23 /queen mounts against 22, and an unguarded mount nobody allowlisted. The route is a server-to-server door for the app render proxy and has its own guard: a bearer equal to QUEEN_CONTRIBUTOR_PROXY_TOKEN (32+ bytes, timingSafeEqual) plus a verified contributor header, and it is off while that token is unset. The trusted-origin check would refuse its only caller, so it is allowlisted with that reason and the pins are re-measured. No other number moved. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2 * fix(queen): keep the earnings status literal in earningsOfLogin The base object literal widened EARNINGS_STATUS to string, so the function no longer matched EarningsOfLogin. Typing base as Omit<EarningsOfLogin, 'earnings'> keeps the literal. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Retargets the runner cabinet at the branch production builds from. Conflicts, all additive: - server.ts: runner mounts beside /queen/public-credits and /queen/public-earnings. - pg-migrate.ts: queen_runner beside queen_tri_earnings. - route-guard-audit.mjs: the runner allowlist entries beside the production branch's /queen/contributor-keys entry (same reason text). - route-guard.test.ts: the production branch's pins plus the two runner mounts: 51 mounts; 28 /queen mounts as 10 public-read, 9 wrapper, 9 allowlisted; eleven unguarded without the allowlist. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
|
I've retargeted this PR to The conflicts were all cases where both sides added something:
The diff against the new base is the runner cabinet alone, 10 files. The #520 CI fixes are already on the base. Local results: #521 has been retargeted to the same branch. It contains this PR, so merge this one first. Generated by Claude Code |
The setup lines fetched queen-runner.mjs from feat/queen-supervisor, where it does not exist (a 404, as the review found). The runner PRs (gHashTag/BrowserOS#518, #521) now target fix/queen-worker-provider-and-prompt-size, the branch the deployed Queen is built from, so the script and its README are linked there, through one RUNNER_BRANCH constant. The contract's "no provider key" check matched the bare word "provider", which that branch name contains. It now matches a key in any spelling (API_KEY, sk-, provider key) and proves it still catches each of them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
|
The Here is the evidence for this head in the meantime:
A maintainer can start the run with Re-run or Run workflow ( Generated by Claude Code |
…has it The MY RUNNERS panel is live on app.t27.ai, and the server half (gHashTag/BrowserOS#518) is not deployed yet: the production Queen answers /queen/me/runners with 404. The panel read every non-OK answer as an outage and told each signed-in person "The Queen did not answer. Try again in a minute." while she was answering. A server that has the cabinet answers that path 200, 401 or 503, never 404, so a 404 on list or create now means "not here yet": the panel says runners are on their way and switches on by itself once the server offers them. No create form is shown in that state. 503 and network failures still read "did not answer". Measured before the change: the production server answers the path 404 with Access-Control-Allow-Origin: https://app.t27.ai, so the browser does see the status. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
What and why
The leaderboard asks people to "lend the swarm a key", but most providers forbid handing a key to a third party (see
how-to-join-the-swarm). The design that asks nobody to breach their terms is the one where the key never moves: the runner runs on the lender's machine, under the lender's account.This PR is the registry half of that design.
/queen/me/runners(GET/POST/DELETE): the cabinet for a person signed in on app.t27.ai. The server sends the person's bearer to the issuer (vibee-render/mcpwhoami) and gets back theirtelegram_id./queen/runner/heartbeat(POST, runner token): marks the runner online and returns its lane. It answerswork: nulland says honestly that handing out tasks is the next stage.queen_runnertable inMIGRATION_SQL. A runner's lane is100000000 + id, akey_indexblock that no operator pool reaches. Dispatch rows on that lane are therefore credited by the existing leaderboard arithmetic.telegram_id), never by name. A runner cannot merge into an operator's row, and it never gets a link to an unverified GitHub login./queen/me/*: onlyhttps://app.t27.ai,Authorization+Content-Type, no credentials. It is mounted beforetrustedCorsMiddleware, like the public routes./queen. No other count changed.No route accepts, stores or returns a provider key.
Website counterpart: gHashTag/trinity#1205 — the MY RUNNERS panel on the LEADERBOARD tab. Deploy this PR first.
Verification
bun run test:api: 1315 pass, 0 fail, run after rebasing onto6bf7efe(feat(queen): public board carries the Queen's verdict and judged head #517).tests/api/queen-runners.test.ts(22 tests) covers:whoami, and telling a refused token (401) apart from an unreachable issuer (503);bun run typecheckis clean, andbiome checkis clean on the changed files.bun run test:pglivewithTRIOS_PG_TEST_URLpointing at a local Postgres 16: 3 pass.limit), the heartbeat by hash, and revoke, including that someone else's revoke returns false;100000002appeared on the leaderboard as{name: "Alice", runner: true, xp: 320}.Limitations / next stage
queen-<issue>inside the container.reapDispatchesFromPreviousBootreaps every unfinished row, so remote rows need to be excluded.whoamifor verification. It is never logged or stored; the cache is keyed by its SHA-256. A narrowly scoped token from vibee-render would be better.queen_runnerat boot.TRIOS_APP_IDENTITY_URLis optional; the default is the production vibee-render.🤖 Generated with Claude Code
https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2