chore(agent): trim mcp-manager dead exports and unused devDeps - #1176
Conversation
Drives the browseros-agent fallow report from 17 issues down to 4 (the residual 4 are all in the remote-hermes surface added by #1174 and out of scope here). mcp-manager: - Slim the public barrel to the 7 symbols real consumers import (routes, main, the two reconcile/service test files). Drop the re-exports of BROWSEROS_MCP_SERVER_NAME, BROWSEROS_MCP_STDIO_SERVER_NAME, getMcpManager, ReconcileUrlInput, InstallAgentResult, McpAgentId, McpAgentRow, ReconcileResult, UninstallAgentResult — all of which are only used by sibling files inside mcp-manager/ itself. - Narrow BROWSEROS_SERVER_NAMES and planFor in service.ts to module-private (used only inside service.ts). This also resolves the AgentServerPlan private-type-leak since planFor becomes private. - Drop the dangling McpAgentIdentifier type alias (zero consumers). UI: - Export IntegrationsSectionProps so the exported IntegrationsSection no longer references a private type. - Narrow AGENT_PRESENTATION in integrations-section.helpers.ts to module-private (used only by presentationFor() in the same file). Deps: - Remove dotenv and picocolors from devDependencies — neither is imported anywhere in the package (dotenv is even explicitly documented as not needed in the agent README).
✅ Tests passed — 1212/1218
|
Greptile SummaryThis PR trims dead exports from the
Confidence Score: 5/5Safe to merge — all dropped exports are confirmed to have zero external consumers, and all narrowed/added visibility changes are strictly non-breaking. Every removed barrel export was cross-checked against all import sites across the workspace; the routes file, main.ts, and test files all import only symbols that remain in the barrel. Narrowing BROWSEROS_SERVER_NAMES and planFor to module-private is a visibility reduction with no call-site impact. Exporting the three props interfaces is purely additive. The devDependency removals are backed by the lock-file diff showing dotenv correctly re-scoped to transitive-only slots. No files require special attention. Important Files Changed
Reviews (2): Last reviewed commit: "chore(agent): restore dotenv/picocolors,..." | Re-trigger Greptile |
…s, gate CI
Restoring the two devDeps that were wrongly dropped (build-script tests
on CI failed: `Cannot find package 'picocolors' / 'dotenv'`). Both are
actually consumed by `scripts/build/{server,cli}.ts` and `scripts/build/log.ts`,
but those entry points are outside fallow's discovery (the package.json
script paths use parent-directory traversal that fallow skips). Listing
both in `.fallowrc.json` `ignoreDependencies` matches the existing
convention used for `pino-pretty`.
Fixes the 4 remaining fallow findings introduced by #1174 so the new CI
gate can be green from day one:
- Export `RemoteHermesBootPillProps` (purely additive — resolves the
private-type-leak on `RemoteHermesBootPill`, and incidentally surfaces
the embedded `RemoteHermesVmStatus` as part of its public shape).
- Export `SocketState` in `ws-bridge.ts` (purely additive — resolves the
private-type-leak on the diagnostic-exposed `snapshot()` method).
- Annotate the orphan `PostTurnResult` with `// fallow-ignore-next-line
unused-type` so it stays available for follow-up wiring without
blocking CI.
Adds a `runner / Fallow` job to `.github/workflows/code-quality.yml`
parallel to Biome and Typecheck. Same shape: checkout → setup-bun →
`bun ci` → `bun fallow`. PRs that touch `packages/browseros-agent/**`
now gate on the dead-code report.
* feat(agent): Remote Hermes provider — Cloudflare-managed Fly VM runtime (#1174) * chore: bump .internal-docs submodule * feat(agent): add Remote Hermes provider backed by Cloudflare control plane Adds a new 'remote-hermes' provider that runs the Hermes agent in a managed Fly VM provisioned via the Cloudflare agent-control-worker. Per-install VM identified by browserosId; chat turns proxy through the worker (HTTP + SSE), and the VM dispatches tool calls back to the laptop's local BrowserOS MCP over a single WebSocket held open by apps/server. No API key, base URL or model required at provider-add. Agent UI - New provider type 'remote-hermes' with Sparkles icon, surfaced first in the Settings template grid with the orange "Recommended" treatment. - Add-provider flow needs only a name; backend handles credentials. - Inline boot pill (RemoteHermesBootPill) shows live progress through the cold-start stages (pulling_image, booting, healthchecking). - Delete provider triggers /remote-hermes/destroy for the last entry. apps/server - lib/remote-hermes/: env, HS256 JWT minting (jose), frame parser, partysocket-backed WS bridge with mutex/refcount/idle-close, RPC router dispatching browseros tool calls to 127.0.0.1:<port>/mcp, ProtocolEvent -> AI SDK UIMessageStream translator, and the turn streamer with cold-start polling (180s budget against /vm/status). - /chat forks on provider='remote-hermes' and pipes the worker's SSE through createUIMessageStreamResponse - side panel sees the same AI SDK stream format as any other provider. - New /remote-hermes route: POST /start, POST /destroy, GET /status. Lifecycle endpoints fire-and-forget; status proxies the worker. Plus collateral lint:fix touch-ups in eval/probeAgent/managedBlock. * fix(agent): address remote-hermes review feedback - bridge: close+null any existing socket at the top of doOpen() so a ReconnectingWebSocket that opens AFTER our 5s OPEN_DEADLINE_MS rejected cannot fire 'open' on the bridge's behalf and start a parallel ping/idle-sweep loop on the wrong socket. Every event handler now checks `this.socket === sock` and returns early when the dispatched socket isn't the live one. - turn: derive the cold-start error message and the boot-poll comment from COLD_START_BUDGET_MS instead of the stale literal "90 seconds" left over from the original 90s budget. - frames: drop unused PONG_FRAME export. pong frames are only synthesized by Cloudflare's setWebSocketAutoResponse worker-side; the laptop never emits one. - remote-hermes route: drop the dead `method: 'POST'` parameter from fireVmLifecycle; both call sites passed the literal and the function always POSTs. * refactor(server): rework Remote Hermes layering to match KlavisClient pattern Addresses code-quality feedback on the original PR. Three problems with the previous shape: 1. Per-request env reads. loadRemoteHermesEnv() ran on every /chat fork and every /remote-hermes/* handler. Now AGENT_RUNNER_JWT_SECRET lives in INLINED_ENV (build-time inlined) and the worker URL is EXTERNAL_URLS.AGENT_CONTROL_WORKER. Both read once at module load, never re-read. 2. Mixed responsibilities in lib/remote-hermes/. The flat dump conflated the HTTP wire client, the WS bridge, the SSE pump, JWT minting, env parsing, and the route handler logic. Split following the existing Klavis precedent: lib/clients/remote-hermes/ remote-hermes-client.ts raw HTTP wrapper (mintJwt + fetch) ws-bridge.ts persistent WS, refcount/idle/race-safe auth.ts, frames.ts, rpc-router.ts, event-translator.ts constants.ts module-internal tunables only api/services/remote-hermes/ remote-hermes-service.ts high-level facade. owns bridge lifecycle, exposes warm/teardown/status/streamTurn, no env or fetch lives here 3. Hidden singleton + inline lifecycle logic in handlers. getBridge() was a module-level singleton constructed inside the /chat handler. Now the service is constructed once in createHttpServer() when INLINED_ENV.AGENT_RUNNER_JWT_SECRET is present (warn-only when absent, matches Klavis behaviour), threaded into ChatRouteDeps + the new RemoteHermesRouteDeps, and closed on Application.shutdown. Net effect: /chat fork: 50 lines -> 10 lines /remote-hermes routes: 119 lines -> 47 lines lib/clients/remote-hermes: cleaner per-file responsibility All [remote-hermes] template logs replaced with structured fields `module: 'remote-hermes'`, matching the rest of the codebase. Shared constants moved to packages/shared/src/constants/hermes.ts: REMOTE_HERMES_PROVIDER_TYPE, REMOTE_HERMES_AGENT_KIND, REMOTE_HERMES_DEFAULT_AGENT_ID. EXTERNAL_URLS.AGENT_CONTROL_WORKER added — matches KLAVIS_PROXY shape. No behaviour change: chat turns, boot pill, cold-start poll, WS bridge race fixes from the prior commit, /vm/start /vm/destroy /vm/status all preserved end-to-end. * fix(server): strip duplicated browseros_browseros_ tool name prefix Tool cards in the side panel rendered as "Mcp browseros browseros suggest app connection" instead of "suggest_app_connection". Cause: acpx normalizes the VM catalog's "<server>.<tool>" dot into an underscore when emitting tool names. Combined with the MCP server name we configure ("browseros") and the catalog server name (also "browseros" since the worker fix in this branch), the wire name becomes "browseros_browseros_suggest_app_connection". Our existing strip only handled the double-underscore acpx prefix and the dot-separated catalog prefix. Add a third pattern that matches "<word>_<word>_" only when the two word groups are identical, so we never chew the head off an unrelated tool that happens to start "browseros_". * refactor(remote-hermes): post-review code-quality sweep apps/server: - Drop the now-unused Remote Hermes event translator — the runtime service emits AI SDK UI Message Stream parts directly, so the laptop just forwards them. - warm()/teardown() now throw on non-2xx from the worker so the route's .catch logs a real error instead of swallowing the failure. - pumpEvents() dismisses the boot pill in finally — handles the edge where the stream ends before any non-`start` part arrives. - /chat fork logs a real reason when remote-hermes hits a server with the service unconfigured. apps/agent: - Add REMOTE_HERMES_PROVIDER_TYPE in lib/llm-providers/types.ts (local mirror of @browseros/shared since the WXT bundle doesn't depend on the shared package) and use it in isRemoteHermesType + ProviderTemplatesSection. Replace the double-filter pin pattern with a sort + Fragment-after-Hermes layout so the order is expressed declaratively. * chore(agent): trim mcp-manager dead exports and unused devDeps (#1176) * chore(agent): trim mcp-manager dead exports and unused devDeps Drives the browseros-agent fallow report from 17 issues down to 4 (the residual 4 are all in the remote-hermes surface added by #1174 and out of scope here). mcp-manager: - Slim the public barrel to the 7 symbols real consumers import (routes, main, the two reconcile/service test files). Drop the re-exports of BROWSEROS_MCP_SERVER_NAME, BROWSEROS_MCP_STDIO_SERVER_NAME, getMcpManager, ReconcileUrlInput, InstallAgentResult, McpAgentId, McpAgentRow, ReconcileResult, UninstallAgentResult — all of which are only used by sibling files inside mcp-manager/ itself. - Narrow BROWSEROS_SERVER_NAMES and planFor in service.ts to module-private (used only inside service.ts). This also resolves the AgentServerPlan private-type-leak since planFor becomes private. - Drop the dangling McpAgentIdentifier type alias (zero consumers). UI: - Export IntegrationsSectionProps so the exported IntegrationsSection no longer references a private type. - Narrow AGENT_PRESENTATION in integrations-section.helpers.ts to module-private (used only by presentationFor() in the same file). Deps: - Remove dotenv and picocolors from devDependencies — neither is imported anywhere in the package (dotenv is even explicitly documented as not needed in the agent README). * chore(agent): restore dotenv/picocolors, fix remaining fallow findings, gate CI Restoring the two devDeps that were wrongly dropped (build-script tests on CI failed: `Cannot find package 'picocolors' / 'dotenv'`). Both are actually consumed by `scripts/build/{server,cli}.ts` and `scripts/build/log.ts`, but those entry points are outside fallow's discovery (the package.json script paths use parent-directory traversal that fallow skips). Listing both in `.fallowrc.json` `ignoreDependencies` matches the existing convention used for `pino-pretty`. Fixes the 4 remaining fallow findings introduced by #1174 so the new CI gate can be green from day one: - Export `RemoteHermesBootPillProps` (purely additive — resolves the private-type-leak on `RemoteHermesBootPill`, and incidentally surfaces the embedded `RemoteHermesVmStatus` as part of its public shape). - Export `SocketState` in `ws-bridge.ts` (purely additive — resolves the private-type-leak on the diagnostic-exposed `snapshot()` method). - Annotate the orphan `PostTurnResult` with `// fallow-ignore-next-line unused-type` so it stays available for follow-up wiring without blocking CI. Adds a `runner / Fallow` job to `.github/workflows/code-quality.yml` parallel to Biome and Typecheck. Same shape: checkout → setup-bun → `bun ci` → `bun fallow`. PRs that touch `packages/browseros-agent/**` now gate on the dead-code report. * feat(agent-mcp-ui): install AI Elements via shadcn registry 48 components from elements.ai-sdk.dev pulled in via: bunx shadcn@latest add @ai-elements/<each component> Catalog covered: Chatbot attachments, chain-of-thought, checkpoint, confirmation, context, conversation, inline-citation, message, model-selector, plan, prompt-input, queue, reasoning, shimmer, sources, suggestion, task, tool Code agent, artifact, code-block, commit, environment-variables, file-tree, jsx-preview, package-info, sandbox, schema-display, snippet, stack-trace, terminal, test-results, web-preview Voice audio-player, mic-selector, persona, speech-input, transcription, voice-selector Workflow canvas, connection, controls, edge, node, panel, toolbar Utilities image, open-in-chat Files land in components/ai-elements/ and consume shadcn primitives from components/ui/, which the same install grew from 3 to 25 to cover all the dependencies (accordion, command, dialog, dropdown, hover-card, popover, scroll-area, select, tabs, tooltip, etc.). Peer deps brought in (devDependencies left untouched; runtime only): ai, streamdown, @streamdown/{cjk,code,math,mermaid}, shiki, @xyflow/react, motion, @rive-app/react-webgl2, media-chrome, cmdk, embla-carousel-react, lucide-react, nanoid, tokenlens, use-stick-to-bottom, react-jsx-parser, ansi-to-react, @radix-ui/react-use-controllable-state. @base-ui/react bumped to ^1.5.0 because the components target a newer API surface than the previous ^1.0.0-beta.6 exposed. Six ai-elements files (attachments, context, inline-citation, plan, prompt-input, voice-selector) ship with `closeDelay`/`openDelay`/ event-handler shapes that don't typecheck against @base-ui/react 1.5.0. Bundled JS runs fine (unknown props get spread onto DOM and silently ignored, event handler arities are forgiving at runtime), only tsc strictness flags them. Marked with `// @ts-nocheck` at the top of each, with a comment explaining the posture. biome.json now disables both formatter and linter for components/ui/** and components/ai-elements/** (third-party drop-ins; both shadcn-installed quote style and the AI Elements internal patterns differ from the repo style). organizeImports assist action also off on the same paths. Verified: bun run typecheck clean bunx biome check clean (89 files) bunx wxt build clean, chrome-mv3 397 kB * feat(agent-mcp-ui): cockpit shell + sidebar + 4 routes Foundation pass on top of the WXT bootstrap. Sidebar with hover-expand behaviour matching apps/agent's idiom, four routed surfaces (cockpit, agents, governance, mcp), plus a stub /agents/new for the future wizard. The cockpit page is the full dashboard design from the prototype, hooked up to mock react-query-kit hooks shaped so the eventual swap to real agent-mcp-interface routes is a fetcher body change. Layout entrypoints/app/App.tsx HashRouter with single layout route (CockpitShell) wrapping 5 children entrypoints/app/main.tsx TooltipProvider added, @fontsource imports for Schibsted Grotesk + Newsreader italic + JetBrains Mono components/layout/CockpitShell.tsx fixed sidebar (w-14 collapsed, w-64 expanded, 150ms collapse delay) + main outlet, pl-14 offset components/layout/PlaceholderScreen shared "coming soon" composite Sidebar components/sidebar/AppSidebar.tsx branding + navigation, no user footer until we have a setting to surface components/sidebar/SidebarBranding.tsx orange B mark + wordmark components/sidebar/SidebarNavigation.tsx 4 lucide-iconed NavLinks with base-ui Tooltip (render prop, not asChild) on collapsed Cockpit surfaces components/cockpit/CockpitHero.tsx hero with serif italic accent components/cockpit/WaitingStrip.tsx container for approvals + handoffs components/cockpit/ApprovalBanner.tsx 3-button approval card components/cockpit/HandoffRow.tsx amber "take over" row components/cockpit/RunningGrid.tsx auto-fill grid + live count chip + AddAgentTile at the end components/cockpit/RunningCard.tsx mini-screencast + label + status + task + watch/stop components/cockpit/AddAgentTile.tsx dashed-border "+ New profile" tile linking to /agents/new components/cockpit/RecentActivity.tsx list container with flagged- count chip components/cockpit/ActivityRow.tsx per-row status icon + agent dot + jump-to action components/cockpit/StatusBadge.tsx token-driven status pill components/cockpit/MiniScreencast.tsx placeholder card-top tile Placeholder screens screens/cockpit/Cockpit.tsx rewrite: composes the surfaces above against mock hooks screens/agents/Agents.tsx placeholder screens/governance/Governance.tsx placeholder screens/mcp/Mcp.tsx placeholder screens/new-agent/NewAgent.tsx placeholder Data modules/api/agents.hooks.ts useAgents (mock; 3 running rows) modules/api/waiting.hooks.ts useApprovals + useHandoffs (1 + 1) modules/api/activity.hooks.ts useRecentActivity (4 rows: blocked, needs-human, allowed, done) lib/status.ts RunStatus union + STATUS_META map + isActiveStatus / isEndedStatus helpers; single source of truth so colors stay consistent across the cockpit, audit, and activity log Design tokens entrypoints/app/styles.css full BrowserOS warm-cream palette wired via @theme inline (shadcn primitives re-pointed at BrowserOS surfaces; bespoke ink scale, status palette, accent ink, shadows, pulse-dot / fade keyframes). Body carries the design's layered radial gradient. Selection uses accent tint instead of chrome blue. Deps @fontsource-variable/schibsted-grotesk @fontsource-variable/jetbrains-mono @fontsource/newsreader (400-italic + 500-italic only) Verified: bun run --filter @browseros/agent-mcp-ui typecheck clean bunx biome check clean (113 files) bunx wxt build clean (880 kB: 317 kB JS, 80 kB CSS, ~480 kB fonts) * fix(agent-mcp-ui): declare the browserOS permission Without `browserOS` in the manifest's permissions array, BrowserOS Chromium's new-tab override gate refuses the extension's claim and the cockpit never replaces chrome://newtab. The apps/agent extension sits on the same permission for the same reason. Adds two near-neighbours that the cockpit will reach for soon: webNavigation for routing the future live-run jump targets (the rest were already declared) * fix(agent-mcp-ui): switch to WXT's conventional newtab entrypoint Renames entrypoints/app/ to entrypoints/newtab/ so WXT auto-wires manifest.chrome_url_overrides.newtab against the generated newtab.html. Drops the hand-rolled chrome_url_overrides block from wxt.config.ts since WXT now manages it. The output file is newtab.html instead of app.html; nothing internal references the page by filename so no other code change is needed. Build output verified: manifest carries { chrome_url_overrides: { newtab: 'newtab.html' } } automatically. Reference: https://wxt.dev/guide/essentials/entrypoints.html#newtab * chore(agent-mcp-ui): wire react-doctor into lint Lint now runs biome and react-doctor concurrently as a single pass via concurrently --group, so findings from both tools land in one terminal output and the combined exit code is non-zero if either fails. react-doctor is invoked through bunx (no devDep yet) so the bun release-age gate doesn't block its newest versions. Config in doctor.config.json mirrors biome's vendored-path ignores (components/ui, components/ai-elements, build outputs). * chore(agent-mcp-ui): add .gitignore and untrack .wxt artifacts Mirrors apps/agent's .gitignore. .wxt/ is regenerated on every wxt dev/build, so the seven previously-tracked files in it churned the diff for no reason. * chore: added verbose flag * chore(agent-mcp-ui): split lint and react-doctor into separate scripts Running them together wasn't worth the extra plumbing. lint stays biome-only; react-doctor moves to a dedicated lint:doctor script with --verbose. Drops concurrently from devDeps. * feat(agent-mcp-ui): new-agent wizard at /agents/new Replaces the placeholder with a 4-section wizard (harness, logins, tool approvals, ACL rules) plus a sticky preview rail showing the MCP URL and an Add-to-harness CTA. Form wires through react-hook-form with a zod schema and shadcn's Form primitive; the submit fires useCreateAgent (react-query-kit mutation, mocked latency) and flips the rail into an added state with a Done button that returns to /agents. Adds shadcn form/label/radio-group/toggle/toggle-group primitives. form.tsx is hand-written for the base-vega Label since the registry copy depends on @radix-ui/react-label which the project doesn't ship. Approvals row uses ToggleGroup with single-select; ACL rows split into a toggle button + sibling trash button to clear the nested-interactive lint warning. doctor.config.json sets deadCode:false because react-doctor's unused-file rule cannot trace WXT's entry resolution and was flagging every shipped component. * feat(agent-mcp-ui): governance hub + audit tab Replaces the placeholder at /governance with a 4-tab hub (Audit, Permissions, Site Rules, Grants) backed by nested routes. The shell renders a sticky header with a pulse-dot live-run counter and a shadcn Tabs nav whose triggers navigate per-tab URLs; the matched sub-route renders inside the Outlet. Audit tab ships: filter chips (All/Running/Blocked/Completed) via ToggleGroup, run-count line, and a list of AuditRow cards (status icon + agent/harness + status pill + meta) that link out to /governance/audit/:runId/replay when clicked. Runs come from a new useRuns mock that mirrors the eventual hono-rpc shape. Permissions, Site Rules, and Grants render a small ComingSoonTab stub so the tab nav feels complete and the URL space is reserved. * fix(agent-mcp-ui): tab + chip active states Base-ui Toggle emits aria-pressed, not data-state=on / data-pressed=true, so the governance filter chips and the new-agent approval toggles weren't actually flipping color when selected. Swapping the selectors to aria-pressed: ties the visual state to the primitive's real attribute. Active governance tab now shows an accent-orange underline. The base shadcn tabs.tsx applies after:bg-foreground and after:bottom-[-5px] in its own className, and Tailwind utility cascade ordering meant my override classes were being beaten. Adding ! to after:bg-accent and after:bottom-[-1px] wins the cascade and lands the underline right on the TabsList's bottom border. * fix(agent-mcp-ui): tabs primitive matches base-ui's data-orientation Base-ui's Tabs root emits data-orientation=horizontal but the shadcn base-vega tabs.tsx was selecting on data-horizontal (no attribute value), which never matched. With the gate failing, the active-tab underline's after:inset-x-0 and after:h-0.5 were dropped, leaving the pseudo-element at 0x0 and invisible. Side-by-side with the prototype made the gap obvious: same accent color and bottom offset, but no underline rendered. Swapping the four group-data-horizontal/group-data-vertical selectors to group-data-[orientation=horizontal] / group-data-[orientation=vertical] lets them match what base-ui actually emits; the active trigger now paints the 2px accent underline at the TabsList border, matching the prototype. * feat(agent-mcp-ui): live-run view with approval and handoff overlays New full-bleed route /run/:runId sits outside the CockpitShell wrapper and gives every agent run its own watch view: a stubbed browser viewport on the left with fake chrome, a centred site host placeholder, the persistent agent-driving badge, and a working pill spelling out the live action; a docked activity panel on the right with the action log, a pinned approval card when the run needs an OK, a pinned handoff notice when it needs the user, an inline block notice when Site Rules killed an action, plus elapsed/tokens/steps stats and pause/stop controls. Approval card honours the v1 UX spec's three-button shape (Allow once, Always allow on domain, Block) and the scope sentence that pins the permission to the current domain. Handoff banner is a full overlay over the viewport with the amber top strip, dimmed page, and an I'm-not-a-robot challenge stub standing in for the real site's CAPTCHA/2FA; the matching in-panel HandoffNotice means the user can resume from either surface. Local state dismisses approvals, handoff, and block notices since the backend isn't wired yet. Cockpit RunningCard 'Watch' button now navigates to /run/<agentId>; run fixtures key by agent id so the cockpit-to-live flow lands on real data for the three running agents. Mock useRun mirrors the eventual hono-rpc /runs/:id + SSE shape. * ci: run code-quality on PRs targeting feat/agent-mcp-interface-bootstrap Stacked PRs on the agent-mcp-interface bootstrap branch were skipping biome / typecheck / fallow because the workflow's pull_request filter only matched main and dev. Adding the explicit branch lets the quality gate fire on every stacked PR before it lands on the parent. --------- Co-authored-by: shivammittal274 <56757235+shivammittal274@users.noreply.github.com>
…P attach, tab attribution (#1333) * feat(agent-mcp-interface): bootstrap package skeleton First slice of the new BrowserOS v2 backend, per the architecture plan. packages/browseros-agent/apps/agent-mcp-interface/ package.json hono dep only; @browseros/agent-mcp-interface tsconfig.json extends monorepo root, composite + emit decls biome.json extends "//", noConsole/noProcessEnv on by default src/ shared/port.ts PROD_API_PORT = 9200 (distinct from 9000/9100/9300) lib/logger.ts Structured JSON to stderr, pino-shaped fields lib/errors.ts HttpError; { error: string } JSON shape env.ts Single chokepoint for process.env reads local-server-url.ts Write-once-after-bind module singleton routes/system.ts /system/health, /system/version, /system/url server.ts Chained .route('/') composition; exports type AppType = typeof routes for the future agent-mcp-ui hono-rpc client main.ts Bun.serve on 127.0.0.1:PROD_API_PORT; sets localServerUrl + logs the bound URL Wiring: packages/browseros-agent/tsconfig.json references += this package packages/browseros-agent/.fallowrc.json entry += src/main.ts Verified: bun run --filter @browseros/agent-mcp-interface typecheck clean bunx biome check apps/agent-mcp-interface/ clean bunx fallow check no new findings bun src/main.ts + curl /system/{health,version,url} ok byok ai-sdk path and the existing apps/server / apps/agent packages are not touched. * chore: minor edits * chore(agent-mcp-interface): scope fallow config to the package Pulls the agent-mcp-interface entry back out of the monorepo-wide packages/browseros-agent/.fallowrc.json and gives the package its own .fallowrc.json. Running `bun run fallow` from inside the package now analyses just this package against its single entry (src/main.ts); running fallow from the agent root no longer drags this package in. Also adds a `fallow` script to the package.json so the command is reachable via `bun run --filter @browseros/agent-mcp-interface fallow`. * Revert "chore(agent-mcp-interface): scope fallow config to the package" This reverts commit 02390dca173fbb9253534b1868fa64f6540a6286. * feat(agent-mcp-ui): bootstrap WXT extension surface Second package of the BrowserOS v2 split. WXT extension that targets the new-tab page and (eventually) the side panel + content overlays. First slice proves the full pipe: React + shadcn (base-vega) renders, TanStack Router resolves routes, TanStack Query + react-query-kit hooks fetch through a typed hono-rpc client, and the runtime closes the loop against the agent-mcp-interface server bound on the same machine. packages/browseros-agent/apps/agent-mcp-ui/ package.json WXT, React 19, Tailwind 4, TanStack Router 1.x, TanStack Query 5, react-query-kit, @base-ui/react, @tabler/icons-react, @browseros/agent-mcp-interface workspace dep (type-only AppType + the PROD_API_PORT constant) tsconfig.json extends .wxt/tsconfig.json, jsx react-jsx, types chrome + bun, paths @/* -> ./* biome.json extends "//", enables tailwindDirectives, ignores routeTree.gen.ts / .wxt / .output / dist, relaxes lint inside components/ui + ai-elements wxt.config.ts @wxt-dev/module-react, manifest with chrome_url_overrides.newtab + 5 permissions, Vite plugins: tanstackRouter (enforce: 'pre', autoCodeSplitting off for now) + tailwindcss components.json shadcn base-vega style, neutral, tabler icons entrypoints/app/ index.html + main.tsx (QueryClientProvider + RouterProvider) + styles.css (Tailwind 4 @theme inline tokens) lib/utils.ts cn() helper (clsx + tailwind-merge) components/ui/ button.tsx, card.tsx, badge.tsx via shadcn CLI modules/api/ client.ts hc<AppType>(baseUrl) + lazy Proxy that re- resolves the URL on each property access queryClient.ts retry: 1, staleTime: 30_000 parseResponse.ts ApiError thrower with .status + .body system.hooks.ts First react-query-kit hooks against the /system/{health,version,url} endpoints routes/ __root.tsx + index.tsx (file-based routing) routeTree.gen.ts Generated by tanstackRouter plugin screens/cockpit/ Minimal Cockpit.tsx that calls useSystemHealth + useSystemVersion, renders shadcn Card + Badge, shows the interface server's name + version on success and a copy-pasteable hint on failure Also exposes shared/port from @browseros/agent-mcp-interface as a real runtime export so the UI can dial in to PROD_API_PORT. The server type export stays type-only ("default": null). Verified: bun run --filter @browseros/agent-mcp-interface typecheck clean bun run --filter @browseros/agent-mcp-ui typecheck clean bunx biome check (both packages) clean bunx wxt build (chrome-mv3 production) clean bun src/main.ts + curl /system/{health,version} ok * refactor(agent-mcp-ui): swap TanStack Router for react-router v7 The existing apps/agent extension routes with react-router v7 + HashRouter. agent-mcp-ui was deviating from that for no real reason — TanStack Router's typed loaders aren't exercised at the bootstrap stage, and the plugin-order collision with WXT's module-react had already forced autoCodeSplitting off. Matching the in-repo precedent removes the codegen step, drops two deps, simplifies wxt.config.ts, and shaves ~48 kB off the production bundle (384 → 336 kB). Changes: package.json -@tanstack/react-router, -@tanstack/router-plugin +react-router ^7.12.0 wxt.config.ts Drop prependRouterPlugin helper + the enforce:'pre' workaround. Plugins now: [tailwindcss()]. Module-react injects @vitejs/plugin-react as before. routes/ Deleted. __root.tsx, index.tsx, routeTree.gen.ts gone. entrypoints/app/App.tsx New. Code-based HashRouter + Routes + Route, matching apps/agent's pattern. entrypoints/app/main.tsx Drops createRouter + RouterProvider + module augmentation; renders <App /> inside QueryClientProvider + StrictMode. biome.json Drops the !routeTree.gen.ts ignore. Verified: bun install clean bun run --filter @browseros/agent-mcp-ui typecheck clean bunx biome check clean bunx wxt build clean, 336 kB bun src/main.ts + curl /system/{health,version} ok * feat(agent-mcp-ui): wire wxt dev launch into BrowserOS Chromium Mirrors the existing apps/agent launch shape so `bun run dev` boots BrowserOS with this extension installed instead of stock Chromium. web-ext.config.ts defineWebExtConfig with BrowserOS binary path, dev-sane chromiumArgs (--use-mock-keychain, --disable-browseros-server, --disable-browseros-extensions, --browseros-dock-icon=dev), and a worktree+package-scoped Chromium profile under /tmp/browseros-dev-<worktree>-<packageHash>. Profile dir distinct from apps/agent's so the two dev runs never share state. BROWSEROS_CDP_PORT / BROWSEROS_SERVER_PORT / BROWSEROS_EXTENSION_PORT / BROWSEROS_USER_DATA_DIR / BROWSEROS_BINARY env overrides supported with the same names the agent extension already uses. .env.example Documents the env vars; copy to .env.development to enable them. Every entry is optional. package.json dev → bun --env-file=.env.development wxt build:dev → same, for the development-mode build Verified: bun --print 'await import("./web-ext.config.ts").then(m => m.default)' resolves the BrowserOS binary path + the per-worktree+package profile dir bunx wxt prepare clean bun run typecheck clean bunx biome check clean bunx wxt build clean, 336 kB * feat(agent-mcp-ui): cockpit + new-agent + governance + live-run (#1193) * feat(agent): Remote Hermes provider — Cloudflare-managed Fly VM runtime (#1174) * chore: bump .internal-docs submodule * feat(agent): add Remote Hermes provider backed by Cloudflare control plane Adds a new 'remote-hermes' provider that runs the Hermes agent in a managed Fly VM provisioned via the Cloudflare agent-control-worker. Per-install VM identified by browserosId; chat turns proxy through the worker (HTTP + SSE), and the VM dispatches tool calls back to the laptop's local BrowserOS MCP over a single WebSocket held open by apps/server. No API key, base URL or model required at provider-add. Agent UI - New provider type 'remote-hermes' with Sparkles icon, surfaced first in the Settings template grid with the orange "Recommended" treatment. - Add-provider flow needs only a name; backend handles credentials. - Inline boot pill (RemoteHermesBootPill) shows live progress through the cold-start stages (pulling_image, booting, healthchecking). - Delete provider triggers /remote-hermes/destroy for the last entry. apps/server - lib/remote-hermes/: env, HS256 JWT minting (jose), frame parser, partysocket-backed WS bridge with mutex/refcount/idle-close, RPC router dispatching browseros tool calls to 127.0.0.1:<port>/mcp, ProtocolEvent -> AI SDK UIMessageStream translator, and the turn streamer with cold-start polling (180s budget against /vm/status). - /chat forks on provider='remote-hermes' and pipes the worker's SSE through createUIMessageStreamResponse - side panel sees the same AI SDK stream format as any other provider. - New /remote-hermes route: POST /start, POST /destroy, GET /status. Lifecycle endpoints fire-and-forget; status proxies the worker. Plus collateral lint:fix touch-ups in eval/probeAgent/managedBlock. * fix(agent): address remote-hermes review feedback - bridge: close+null any existing socket at the top of doOpen() so a ReconnectingWebSocket that opens AFTER our 5s OPEN_DEADLINE_MS rejected cannot fire 'open' on the bridge's behalf and start a parallel ping/idle-sweep loop on the wrong socket. Every event handler now checks `this.socket === sock` and returns early when the dispatched socket isn't the live one. - turn: derive the cold-start error message and the boot-poll comment from COLD_START_BUDGET_MS instead of the stale literal "90 seconds" left over from the original 90s budget. - frames: drop unused PONG_FRAME export. pong frames are only synthesized by Cloudflare's setWebSocketAutoResponse worker-side; the laptop never emits one. - remote-hermes route: drop the dead `method: 'POST'` parameter from fireVmLifecycle; both call sites passed the literal and the function always POSTs. * refactor(server): rework Remote Hermes layering to match KlavisClient pattern Addresses code-quality feedback on the original PR. Three problems with the previous shape: 1. Per-request env reads. loadRemoteHermesEnv() ran on every /chat fork and every /remote-hermes/* handler. Now AGENT_RUNNER_JWT_SECRET lives in INLINED_ENV (build-time inlined) and the worker URL is EXTERNAL_URLS.AGENT_CONTROL_WORKER. Both read once at module load, never re-read. 2. Mixed responsibilities in lib/remote-hermes/. The flat dump conflated the HTTP wire client, the WS bridge, the SSE pump, JWT minting, env parsing, and the route handler logic. Split following the existing Klavis precedent: lib/clients/remote-hermes/ remote-hermes-client.ts raw HTTP wrapper (mintJwt + fetch) ws-bridge.ts persistent WS, refcount/idle/race-safe auth.ts, frames.ts, rpc-router.ts, event-translator.ts constants.ts module-internal tunables only api/services/remote-hermes/ remote-hermes-service.ts high-level facade. owns bridge lifecycle, exposes warm/teardown/status/streamTurn, no env or fetch lives here 3. Hidden singleton + inline lifecycle logic in handlers. getBridge() was a module-level singleton constructed inside the /chat handler. Now the service is constructed once in createHttpServer() when INLINED_ENV.AGENT_RUNNER_JWT_SECRET is present (warn-only when absent, matches Klavis behaviour), threaded into ChatRouteDeps + the new RemoteHermesRouteDeps, and closed on Application.shutdown. Net effect: /chat fork: 50 lines -> 10 lines /remote-hermes routes: 119 lines -> 47 lines lib/clients/remote-hermes: cleaner per-file responsibility All [remote-hermes] template logs replaced with structured fields `module: 'remote-hermes'`, matching the rest of the codebase. Shared constants moved to packages/shared/src/constants/hermes.ts: REMOTE_HERMES_PROVIDER_TYPE, REMOTE_HERMES_AGENT_KIND, REMOTE_HERMES_DEFAULT_AGENT_ID. EXTERNAL_URLS.AGENT_CONTROL_WORKER added — matches KLAVIS_PROXY shape. No behaviour change: chat turns, boot pill, cold-start poll, WS bridge race fixes from the prior commit, /vm/start /vm/destroy /vm/status all preserved end-to-end. * fix(server): strip duplicated browseros_browseros_ tool name prefix Tool cards in the side panel rendered as "Mcp browseros browseros suggest app connection" instead of "suggest_app_connection". Cause: acpx normalizes the VM catalog's "<server>.<tool>" dot into an underscore when emitting tool names. Combined with the MCP server name we configure ("browseros") and the catalog server name (also "browseros" since the worker fix in this branch), the wire name becomes "browseros_browseros_suggest_app_connection". Our existing strip only handled the double-underscore acpx prefix and the dot-separated catalog prefix. Add a third pattern that matches "<word>_<word>_" only when the two word groups are identical, so we never chew the head off an unrelated tool that happens to start "browseros_". * refactor(remote-hermes): post-review code-quality sweep apps/server: - Drop the now-unused Remote Hermes event translator — the runtime service emits AI SDK UI Message Stream parts directly, so the laptop just forwards them. - warm()/teardown() now throw on non-2xx from the worker so the route's .catch logs a real error instead of swallowing the failure. - pumpEvents() dismisses the boot pill in finally — handles the edge where the stream ends before any non-`start` part arrives. - /chat fork logs a real reason when remote-hermes hits a server with the service unconfigured. apps/agent: - Add REMOTE_HERMES_PROVIDER_TYPE in lib/llm-providers/types.ts (local mirror of @browseros/shared since the WXT bundle doesn't depend on the shared package) and use it in isRemoteHermesType + ProviderTemplatesSection. Replace the double-filter pin pattern with a sort + Fragment-after-Hermes layout so the order is expressed declaratively. * chore(agent): trim mcp-manager dead exports and unused devDeps (#1176) * chore(agent): trim mcp-manager dead exports and unused devDeps Drives the browseros-agent fallow report from 17 issues down to 4 (the residual 4 are all in the remote-hermes surface added by #1174 and out of scope here). mcp-manager: - Slim the public barrel to the 7 symbols real consumers import (routes, main, the two reconcile/service test files). Drop the re-exports of BROWSEROS_MCP_SERVER_NAME, BROWSEROS_MCP_STDIO_SERVER_NAME, getMcpManager, ReconcileUrlInput, InstallAgentResult, McpAgentId, McpAgentRow, ReconcileResult, UninstallAgentResult — all of which are only used by sibling files inside mcp-manager/ itself. - Narrow BROWSEROS_SERVER_NAMES and planFor in service.ts to module-private (used only inside service.ts). This also resolves the AgentServerPlan private-type-leak since planFor becomes private. - Drop the dangling McpAgentIdentifier type alias (zero consumers). UI: - Export IntegrationsSectionProps so the exported IntegrationsSection no longer references a private type. - Narrow AGENT_PRESENTATION in integrations-section.helpers.ts to module-private (used only by presentationFor() in the same file). Deps: - Remove dotenv and picocolors from devDependencies — neither is imported anywhere in the package (dotenv is even explicitly documented as not needed in the agent README). * chore(agent): restore dotenv/picocolors, fix remaining fallow findings, gate CI Restoring the two devDeps that were wrongly dropped (build-script tests on CI failed: `Cannot find package 'picocolors' / 'dotenv'`). Both are actually consumed by `scripts/build/{server,cli}.ts` and `scripts/build/log.ts`, but those entry points are outside fallow's discovery (the package.json script paths use parent-directory traversal that fallow skips). Listing both in `.fallowrc.json` `ignoreDependencies` matches the existing convention used for `pino-pretty`. Fixes the 4 remaining fallow findings introduced by #1174 so the new CI gate can be green from day one: - Export `RemoteHermesBootPillProps` (purely additive — resolves the private-type-leak on `RemoteHermesBootPill`, and incidentally surfaces the embedded `RemoteHermesVmStatus` as part of its public shape). - Export `SocketState` in `ws-bridge.ts` (purely additive — resolves the private-type-leak on the diagnostic-exposed `snapshot()` method). - Annotate the orphan `PostTurnResult` with `// fallow-ignore-next-line unused-type` so it stays available for follow-up wiring without blocking CI. Adds a `runner / Fallow` job to `.github/workflows/code-quality.yml` parallel to Biome and Typecheck. Same shape: checkout → setup-bun → `bun ci` → `bun fallow`. PRs that touch `packages/browseros-agent/**` now gate on the dead-code report. * feat(agent-mcp-ui): install AI Elements via shadcn registry 48 components from elements.ai-sdk.dev pulled in via: bunx shadcn@latest add @ai-elements/<each component> Catalog covered: Chatbot attachments, chain-of-thought, checkpoint, confirmation, context, conversation, inline-citation, message, model-selector, plan, prompt-input, queue, reasoning, shimmer, sources, suggestion, task, tool Code agent, artifact, code-block, commit, environment-variables, file-tree, jsx-preview, package-info, sandbox, schema-display, snippet, stack-trace, terminal, test-results, web-preview Voice audio-player, mic-selector, persona, speech-input, transcription, voice-selector Workflow canvas, connection, controls, edge, node, panel, toolbar Utilities image, open-in-chat Files land in components/ai-elements/ and consume shadcn primitives from components/ui/, which the same install grew from 3 to 25 to cover all the dependencies (accordion, command, dialog, dropdown, hover-card, popover, scroll-area, select, tabs, tooltip, etc.). Peer deps brought in (devDependencies left untouched; runtime only): ai, streamdown, @streamdown/{cjk,code,math,mermaid}, shiki, @xyflow/react, motion, @rive-app/react-webgl2, media-chrome, cmdk, embla-carousel-react, lucide-react, nanoid, tokenlens, use-stick-to-bottom, react-jsx-parser, ansi-to-react, @radix-ui/react-use-controllable-state. @base-ui/react bumped to ^1.5.0 because the components target a newer API surface than the previous ^1.0.0-beta.6 exposed. Six ai-elements files (attachments, context, inline-citation, plan, prompt-input, voice-selector) ship with `closeDelay`/`openDelay`/ event-handler shapes that don't typecheck against @base-ui/react 1.5.0. Bundled JS runs fine (unknown props get spread onto DOM and silently ignored, event handler arities are forgiving at runtime), only tsc strictness flags them. Marked with `// @ts-nocheck` at the top of each, with a comment explaining the posture. biome.json now disables both formatter and linter for components/ui/** and components/ai-elements/** (third-party drop-ins; both shadcn-installed quote style and the AI Elements internal patterns differ from the repo style). organizeImports assist action also off on the same paths. Verified: bun run typecheck clean bunx biome check clean (89 files) bunx wxt build clean, chrome-mv3 397 kB * feat(agent-mcp-ui): cockpit shell + sidebar + 4 routes Foundation pass on top of the WXT bootstrap. Sidebar with hover-expand behaviour matching apps/agent's idiom, four routed surfaces (cockpit, agents, governance, mcp), plus a stub /agents/new for the future wizard. The cockpit page is the full dashboard design from the prototype, hooked up to mock react-query-kit hooks shaped so the eventual swap to real agent-mcp-interface routes is a fetcher body change. Layout entrypoints/app/App.tsx HashRouter with single layout route (CockpitShell) wrapping 5 children entrypoints/app/main.tsx TooltipProvider added, @fontsource imports for Schibsted Grotesk + Newsreader italic + JetBrains Mono components/layout/CockpitShell.tsx fixed sidebar (w-14 collapsed, w-64 expanded, 150ms collapse delay) + main outlet, pl-14 offset components/layout/PlaceholderScreen shared "coming soon" composite Sidebar components/sidebar/AppSidebar.tsx branding + navigation, no user footer until we have a setting to surface components/sidebar/SidebarBranding.tsx orange B mark + wordmark components/sidebar/SidebarNavigation.tsx 4 lucide-iconed NavLinks with base-ui Tooltip (render prop, not asChild) on collapsed Cockpit surfaces components/cockpit/CockpitHero.tsx hero with serif italic accent components/cockpit/WaitingStrip.tsx container for approvals + handoffs components/cockpit/ApprovalBanner.tsx 3-button approval card components/cockpit/HandoffRow.tsx amber "take over" row components/cockpit/RunningGrid.tsx auto-fill grid + live count chip + AddAgentTile at the end components/cockpit/RunningCard.tsx mini-screencast + label + status + task + watch/stop components/cockpit/AddAgentTile.tsx dashed-border "+ New profile" tile linking to /agents/new components/cockpit/RecentActivity.tsx list container with flagged- count chip components/cockpit/ActivityRow.tsx per-row status icon + agent dot + jump-to action components/cockpit/StatusBadge.tsx token-driven status pill components/cockpit/MiniScreencast.tsx placeholder card-top tile Placeholder screens screens/cockpit/Cockpit.tsx rewrite: composes the surfaces above against mock hooks screens/agents/Agents.tsx placeholder screens/governance/Governance.tsx placeholder screens/mcp/Mcp.tsx placeholder screens/new-agent/NewAgent.tsx placeholder Data modules/api/agents.hooks.ts useAgents (mock; 3 running rows) modules/api/waiting.hooks.ts useApprovals + useHandoffs (1 + 1) modules/api/activity.hooks.ts useRecentActivity (4 rows: blocked, needs-human, allowed, done) lib/status.ts RunStatus union + STATUS_META map + isActiveStatus / isEndedStatus helpers; single source of truth so colors stay consistent across the cockpit, audit, and activity log Design tokens entrypoints/app/styles.css full BrowserOS warm-cream palette wired via @theme inline (shadcn primitives re-pointed at BrowserOS surfaces; bespoke ink scale, status palette, accent ink, shadows, pulse-dot / fade keyframes). Body carries the design's layered radial gradient. Selection uses accent tint instead of chrome blue. Deps @fontsource-variable/schibsted-grotesk @fontsource-variable/jetbrains-mono @fontsource/newsreader (400-italic + 500-italic only) Verified: bun run --filter @browseros/agent-mcp-ui typecheck clean bunx biome check clean (113 files) bunx wxt build clean (880 kB: 317 kB JS, 80 kB CSS, ~480 kB fonts) * fix(agent-mcp-ui): declare the browserOS permission Without `browserOS` in the manifest's permissions array, BrowserOS Chromium's new-tab override gate refuses the extension's claim and the cockpit never replaces chrome://newtab. The apps/agent extension sits on the same permission for the same reason. Adds two near-neighbours that the cockpit will reach for soon: webNavigation for routing the future live-run jump targets (the rest were already declared) * fix(agent-mcp-ui): switch to WXT's conventional newtab entrypoint Renames entrypoints/app/ to entrypoints/newtab/ so WXT auto-wires manifest.chrome_url_overrides.newtab against the generated newtab.html. Drops the hand-rolled chrome_url_overrides block from wxt.config.ts since WXT now manages it. The output file is newtab.html instead of app.html; nothing internal references the page by filename so no other code change is needed. Build output verified: manifest carries { chrome_url_overrides: { newtab: 'newtab.html' } } automatically. Reference: https://wxt.dev/guide/essentials/entrypoints.html#newtab * chore(agent-mcp-ui): wire react-doctor into lint Lint now runs biome and react-doctor concurrently as a single pass via concurrently --group, so findings from both tools land in one terminal output and the combined exit code is non-zero if either fails. react-doctor is invoked through bunx (no devDep yet) so the bun release-age gate doesn't block its newest versions. Config in doctor.config.json mirrors biome's vendored-path ignores (components/ui, components/ai-elements, build outputs). * chore(agent-mcp-ui): add .gitignore and untrack .wxt artifacts Mirrors apps/agent's .gitignore. .wxt/ is regenerated on every wxt dev/build, so the seven previously-tracked files in it churned the diff for no reason. * chore: added verbose flag * chore(agent-mcp-ui): split lint and react-doctor into separate scripts Running them together wasn't worth the extra plumbing. lint stays biome-only; react-doctor moves to a dedicated lint:doctor script with --verbose. Drops concurrently from devDeps. * feat(agent-mcp-ui): new-agent wizard at /agents/new Replaces the placeholder with a 4-section wizard (harness, logins, tool approvals, ACL rules) plus a sticky preview rail showing the MCP URL and an Add-to-harness CTA. Form wires through react-hook-form with a zod schema and shadcn's Form primitive; the submit fires useCreateAgent (react-query-kit mutation, mocked latency) and flips the rail into an added state with a Done button that returns to /agents. Adds shadcn form/label/radio-group/toggle/toggle-group primitives. form.tsx is hand-written for the base-vega Label since the registry copy depends on @radix-ui/react-label which the project doesn't ship. Approvals row uses ToggleGroup with single-select; ACL rows split into a toggle button + sibling trash button to clear the nested-interactive lint warning. doctor.config.json sets deadCode:false because react-doctor's unused-file rule cannot trace WXT's entry resolution and was flagging every shipped component. * feat(agent-mcp-ui): governance hub + audit tab Replaces the placeholder at /governance with a 4-tab hub (Audit, Permissions, Site Rules, Grants) backed by nested routes. The shell renders a sticky header with a pulse-dot live-run counter and a shadcn Tabs nav whose triggers navigate per-tab URLs; the matched sub-route renders inside the Outlet. Audit tab ships: filter chips (All/Running/Blocked/Completed) via ToggleGroup, run-count line, and a list of AuditRow cards (status icon + agent/harness + status pill + meta) that link out to /governance/audit/:runId/replay when clicked. Runs come from a new useRuns mock that mirrors the eventual hono-rpc shape. Permissions, Site Rules, and Grants render a small ComingSoonTab stub so the tab nav feels complete and the URL space is reserved. * fix(agent-mcp-ui): tab + chip active states Base-ui Toggle emits aria-pressed, not data-state=on / data-pressed=true, so the governance filter chips and the new-agent approval toggles weren't actually flipping color when selected. Swapping the selectors to aria-pressed: ties the visual state to the primitive's real attribute. Active governance tab now shows an accent-orange underline. The base shadcn tabs.tsx applies after:bg-foreground and after:bottom-[-5px] in its own className, and Tailwind utility cascade ordering meant my override classes were being beaten. Adding ! to after:bg-accent and after:bottom-[-1px] wins the cascade and lands the underline right on the TabsList's bottom border. * fix(agent-mcp-ui): tabs primitive matches base-ui's data-orientation Base-ui's Tabs root emits data-orientation=horizontal but the shadcn base-vega tabs.tsx was selecting on data-horizontal (no attribute value), which never matched. With the gate failing, the active-tab underline's after:inset-x-0 and after:h-0.5 were dropped, leaving the pseudo-element at 0x0 and invisible. Side-by-side with the prototype made the gap obvious: same accent color and bottom offset, but no underline rendered. Swapping the four group-data-horizontal/group-data-vertical selectors to group-data-[orientation=horizontal] / group-data-[orientation=vertical] lets them match what base-ui actually emits; the active trigger now paints the 2px accent underline at the TabsList border, matching the prototype. * feat(agent-mcp-ui): live-run view with approval and handoff overlays New full-bleed route /run/:runId sits outside the CockpitShell wrapper and gives every agent run its own watch view: a stubbed browser viewport on the left with fake chrome, a centred site host placeholder, the persistent agent-driving badge, and a working pill spelling out the live action; a docked activity panel on the right with the action log, a pinned approval card when the run needs an OK, a pinned handoff notice when it needs the user, an inline block notice when Site Rules killed an action, plus elapsed/tokens/steps stats and pause/stop controls. Approval card honours the v1 UX spec's three-button shape (Allow once, Always allow on domain, Block) and the scope sentence that pins the permission to the current domain. Handoff banner is a full overlay over the viewport with the amber top strip, dimmed page, and an I'm-not-a-robot challenge stub standing in for the real site's CAPTCHA/2FA; the matching in-panel HandoffNotice means the user can resume from either surface. Local state dismisses approvals, handoff, and block notices since the backend isn't wired yet. Cockpit RunningCard 'Watch' button now navigates to /run/<agentId>; run fixtures key by agent id so the cockpit-to-live flow lands on real data for the three running agents. Mock useRun mirrors the eventual hono-rpc /runs/:id + SSE shape. * ci: run code-quality on PRs targeting feat/agent-mcp-interface-bootstrap Stacked PRs on the agent-mcp-interface bootstrap branch were skipping biome / typecheck / fallow because the workflow's pull_request filter only matched main and dev. Adding the explicit branch lets the quality gate fire on every stacked PR before it lands on the parent. --------- Co-authored-by: shivammittal274 <56757235+shivammittal274@users.noreply.github.com> * feat(agent-mcp-ui): replay + agents directory + mcp registry (#1221) * feat(agent-mcp-ui): replay view at /governance/audit/:runId/replay Full-bleed replay player that sits outside CockpitShell so the recorded run gets the whole viewport. The top bar shows the task title, agent and harness, status pill, and a stat strip (duration, tokens, steps, approvals). The body splits into a reconstructed browser viewport on the left (fake chrome + site host placeholder + caption pill that tracks the playhead), a transport with play/pause/restart + native range-input scrubber (overlaid with accent track, kind-coloured bookmark dots for approval/block/done frames, and an accent thumb) + 1x/2x/4x speed toggle, and a right rail Action Timeline whose rows highlight the current frame, dim future frames, and click-to-seek. Playback wallclock lives in usePlayback hook (the project's one allowed useEffect case: starting and cancelling setInterval tied to play state). Scrubber is a real <input type="range"> styled transparently over the visual track so we get native click-to-position, keyboard arrows, Home/End, and screen-reader semantics for free; bookmark buttons sit at z-10 above the track and below the input, so direct clicks still seek to their frame. Mock useReplay keyed by run id mirrors the eventual /runs/:id/replay shape. * feat(agent-mcp-ui): agents directory at /agents with revoke flow Replaces the /agents placeholder with a real directory of configured agent profiles. Header shows the configured-count pill and a primary Add agent CTA that lands the user on the existing /agents/new wizard; the body renders one row per profile with the harness icon chip, name + harness, scope summary (logins, ACL rules, blocked actions, always-allow grants), last-run timestamp, status badge (Configured / Paused / Disabled), and Edit + Revoke buttons. Empty state renders a dashed coming-soon-style card with its own Add CTA. Revoke runs through shadcn AlertDialog (not window.confirm) so focus trapping and ARIA semantics ship for free. useDeleteAgent's onSuccess writes back to the agent-profiles cache via setQueryData so the row vanishes immediately without a refetch, per the project's no-parallel-state-over-cache rule. Edit currently navigates to /agents/:id/edit which is a placeholder slot until the new-agent wizard grows an edit mode. Adds shadcn alert-dialog primitive. Mock useAgentProfiles returns seven profiles spanning every status. * fix(agent-mcp-ui): honour prefers-reduced-motion + snapshot CockpitShell ref Adds a global prefers-reduced-motion: reduce media block in styles.css that collapses every animation and transition to a 0.01ms no-op for users with vestibular sensitivities, satisfying WCAG 2.3.3 across the cockpit (pulse-dot live indicators, sidebar expand, replay scrubber transition, in-app fade-ups). Refactors CockpitShell's unmount cleanup to snapshot the timeout ref object into a stable local before closing over it in the cleanup, which is the React docs' canonical pattern for refs in effects and resolves the missing-effect-dependencies warning. Behaviour is unchanged: the cleanup still clears whatever timeout id is current at unmount time. react-doctor score moves from 74/100 with 2 findings to 100/100 with 0 findings. * feat(agent-mcp-ui): mcp registry at /mcp Replaces the /mcp placeholder with the per-agent MCP endpoint registry. Reads every configured profile from useAgentProfiles and renders one card per profile: harness icon chip + name + harness, slug + CLI hint, status pill, the dark URL block with a copy button, and a Regenerate URL + Add to {harness} button pair. The Add CTA flips into a brief Added confirmation; the copy button flips into a check icon for 1.5s so the user knows the clipboard write landed. Adds useRegenerateMcpUrl mock mutation that rotates the slug and writes the new URL straight back into the agent-profiles cache via setQueryData, so the row reflects the new endpoint without a refetch. Shape mirrors the eventual hono-rpc surface. Skip the per-harness /mcp/setup-* helper screens for now per the running plan; we land them with the onboarding work. * fix(agent-mcp-ui): keep selected text readable on dark surfaces Global ::selection only set a background, so on light-on-dark surfaces like the MCP URL block the cream text disappeared into the accent-tint selection highlight. Pinning the foreground to ink keeps every selection readable: dark ink on light tint everywhere, including the live-run viewport caption and the MCP code block. * chore(browseros-agent): bump biome to 2.5.0 2.5.0 just cleared the bunfig release-age gate so the local install matches what CI's version: latest has already been pulling. Updates the package.json pin plus the schema references in the root biome.json and apps/agent-mcp-ui/biome.json. apps/agent-mcp-ui stays clean under both bun run lint and biome ci. The pre-existing diagnostics surfacing on apps/eval, apps/server, apps/agent, packages/shared, and scripts/dev are unchanged from before the bump. * feat(agent-mcp-ui): onboarding + governance permissions/site-rules/grants + edit wizard (#1223) * feat(agent-mcp-ui): first-launch onboarding flow at /onboarding Adds a four-step onboarding flow sitting full-bleed outside the CockpitShell. Left brand column carries the BrowserOS logo, a Newsreader-italic pull quote, and three value props (fast & token-cheap / logged in as you / under your control). Right column shows step dots up top and one of four step panels: Welcome (set up vs reconnect), Import Logins (Chrome-quit gate, profile picker with default Work + Personal selected, Keychain notice, progress card, summary), Connect to Claude (one-click add or copyable CLI fallback, success card), Ready (two starter prompts with copy buttons, Open BrowserOS CTA). Reconnect and Open BrowserOS both navigate to /. Adds useImportChromeSessions and useConnectToClaude mock mutations (react-query-kit createMutation) whose shape matches the eventual hono-rpc surfaces. CHROME_PROFILES seeds three profiles totalling 55 sites and 14 logins; STARTER_PROMPTS reuses the prompt strings already surfaced elsewhere in the cockpit. Skip first-launch gating for now: per the running plan's open questions, the where-does-the-flag-live decision lives with the backend SSE work. * feat(agent-mcp-ui): governance permissions, site rules, grants tabs Rounds out the governance hub. The three placeholder tabs now ship real surfaces: Permissions: read-only catalog of the six action categories grouped into the three buckets (Auto / Ask / Block) every new agent inherits from. Read straight from new-agent.schemas' APPROVAL_CATEGORIES so the wizard's default verdicts and the catalog can't drift. Layout is a three-column lg grid with verdict-coloured bucket cards. Site Rules: list of (label, domain, action) blocks the browser enforces directly. Each row carries a coloured action badge, the domain in mono, and a delete button. An inline 'Add a rule' form expands into a react-hook-form + zod editor with three fields (label, domain, action select) and submits through useAddSiteRule. setQueryData writes on both add and delete keep the list as the cache's source of truth. Grants: the always-allow ledger. Per-row action + domain + grantee + when + optional note + Revoke button. Revoke routes through a shadcn AlertDialog explaining the consequence (future attempts re-prompt, existing runs unaffected). useRevokeGrant's onSuccess drops the row from the cache. ComingSoonTab is removed; nothing imports it any more. * feat(agent-mcp-ui): edit-mode wizard at /agents/:id/edit + recent-activity replay link Closes the two loose ends called out in the running plan. The new-agent wizard now accepts an optional mode prop ('create' | 'edit'); /agents/:id/edit renders it with mode=edit. In edit mode the data hook reads the agent id from useParams, fetches the wizard-shape values via useAgentProfileDetail (a mock that synthesises full NewAgentValues from an AgentProfile summary), drives the form's reactive values prop, and routes submit through useUpdateAgent. The mutation's onSuccess patches the agent-profiles cache so the directory's row reflects the rename immediately. Header copy, submit CTA, pending label, and success card all flip to edit-mode strings ('Edit agent', 'Save changes to X', 'Saving…', 'X updated'); the copy-from-existing card is hidden in edit mode. Cockpit recent-activity now lands done rows on the replay route. Added an optional runId to the ActivityRow type and a new History-icon Replay button on done rows that links to /governance/audit/:runId/replay. The Codex . Log calls done row points at run-concur-may so the demo lands on real fixture data. * fix(agent-mcp-ui): keep AddSiteRuleForm mounted until mutation settles Previously the form ran close() synchronously after onSubmit, before addRule.mutate could resolve. Today the mock always succeeds; once a real backend lands, a 4xx would silently drop the row and lose the user's input. Widening the onSubmit prop to forward react-query-kit's mutation options lets the parent hand close to the mutation's onSuccess, so the form stays mounted on failure and a FormMessage can surface the error once we wire one in. * feat(agent-mcp-interface): phase 1 — agent profiles CRUD end-to-end (#1224) * chore(agent-mcp-interface): foundation for phase 1 Adds the storage helper Phase 1 leans on, plus the deps and env reads that helper + the upcoming agent routes need. env.ts now also exposes BROWSEROS_DIR overrides and an isDevelopment flag (still the only sanctioned process.env reader). src/lib/browseros-dir.ts resolves <homedir>/.browseros (or .browseros-dev under NODE_ENV=development) with the env override winning; the package writes everything under <browserosDir>/mcp-interface/. src/lib/storage.ts wraps readJson/writeJson/listFiles/removeFile/ensureDir/fileExists around the interface root, validates every read and write with a supplied zod schema, refuses absolute paths or .. escapes, and writes through a <name>.tmp rename so a mid-write crash leaves either prior contents or nothing. Adds bun test wiring + tests/_helpers/temp-browseros-dir.ts so every test gets an isolated tmp root. 13 storage tests pass. * feat(agent-mcp-interface): agent profile schemas + service schemas.ts is the wire contract the UI's typed client picks up via AppType. Mirrors the existing UI wizard shape (NewAgentValues) and adds the storage shape (server-managed id / slug / mcpUrl / status / timestamps) plus the directory projection used by GET / responses. service.ts wraps the storage helper with file-backed CRUD: one profile per file at <browserosDir>/mcp-interface/agents/<id>.json keyed by nanoid(8). Slug is the user-facing identifier and is uniqued across all profiles via uniqueSlug (which collides up to -99 before throwing). mcpUrl is recomputed from getLocalServerUrl on every read so a port change between boots doesn't strand the stored value. lib/slug.ts mirrors the UI's toSlug so wizard preview and persisted slug match. 15 service tests cover create / list / detail / update (rename + slug rotation, slug stability when name unchanged) / remove / regenerate / parallel updates. Plus the original 13 storage tests still pass. 28 tests total in 354ms. * feat(agent-mcp-interface): /agents CRUD routes Thin Hono layer over routes/agents/service.ts: zValidator rejects malformed bodies with structured 400s, missing-id paths surface 404 via HttpError, the rest just translate HTTP shape. Chained into server.ts via .route('/', agentsRoute) so AppType automatically picks up POST /agents, GET /agents, GET /agents/:id, PATCH /agents/:id, DELETE /agents/:id, POST /agents/:id/mcp-url:regenerate. Five route-level integration tests drive the typed client (hc<AppType>) against app.fetch with no real port bind, in an isolated tmp <browserosDir> per case. Covers the full lifecycle, every 404 path, the 400 zod path, slug collision through the route, and parallel updates of two profiles. The regenerate slug regex was tightened to allow nanoid-suffixed multi-hyphen slugs (toSlug normalises any _ in the nanoid output to -). 33 tests / 88 expect calls pass in under 100ms. * feat(agent-mcp-ui): swap six agent hooks for real client calls Replaces the in-memory mocks for useAgentProfiles, useAgentProfileDetail, useCreateAgent, useUpdateAgent, useDeleteAgent, and useRegenerateMcpUrl with hono-rpc calls through the existing client + parseResponse pair. Strips the MOCK_AGENT_PROFILES fixture, the profileToWizardValues synthesiser, the buildMcpUrl/toSlug/nanoid mock helpers, and the artificial setTimeout latencies — the cache surface seen by every consumer (Agents directory, new-agent wizard create + edit, MCP registry regenerate / delete dialogs) stays byte-identical because the wire types now flow from AppType and match what the UI already expected. useAgents (cockpit running grid) stays on its three-row MOCK_AGENTS fixture; that hook becomes Phase 4's projection over the runs store, called out in a top-of-file comment. UI typecheck, lint, lint:doctor all clean. react-doctor holds at 100/100. * fix(agent-mcp-interface): harden phase 1 against the three greptile findings Three independent fixes; reviewer suggestions on PR #1224 covered them all. 1. Storage path guard now inspects the raw input for '..' segments before normalize collapses them. 'agents/../config.json' previously normalized to 'config.json' and slipped past the rooted-prefix check, which would have let any future route forwarding a path-shaped id read or delete files at the mcp-interface/ root. Storage tests cover read/write/remove on a lateral-traversal path. 2. Service layer validates the id shape (matches the nanoid alphabet, length-capped) inside loadById and remove. Traversal-shaped ids on any read/write/delete path now resolve as not-found rather than reaching the storage layer. Service test exercises four evil ids across all four entry points. 3. loadAll uses Promise.allSettled + logger.warn instead of Promise.all so a single corrupt agent json (manual edit, partial migration, half-written file on a weird FS) gets logged + skipped rather than rejecting the whole call. Without this, one bad file would brick list and create until the user manually deleted it. Test writes a garbage file alongside a valid one and confirms list returns only the valid one + create still works. 4. AsyncMutex serialises create / update / regenerateMcpUrl so the read-snapshot → uniqueSlug → write window cannot race against itself. Closes the TOCTOU window where two concurrent same-name creates could both pass the uniqueness check and write the same slug. Reads stay lock-free. Mutex has its own unit tests (FIFO ordering, rejection doesn't block subsequent tasks). Service test fires 10 parallel creates with the same name and asserts 10 distinct slugs come back (race, race-2, ..., race-10). 39 tests / 125 expect calls pass. Lint + typecheck clean. * feat(agent-mcp-interface): site rules + permissions catalog + check api (#1231) * feat(agent-mcp-interface): add domain glob matcher + approval catalog seed * feat(agent-mcp-interface): file-backed site-rules service * feat(agent-mcp-interface): wire /site-rules and /permissions/catalog routes * feat(agent-mcp-interface): permissions.check api for executor pre-flight * feat(agent-mcp-ui): swap site-rules + permissions catalog hooks to real client * fix(agent-mcp-interface): enforce admin site rules + warn on catalog fallback * feat(agent-mcp-interface): per-agent MCP server + executor stub (#1232) * feat(agent-mcp-interface): browser executor interface + deterministic stub * feat(agent-mcp-interface): wire /mcp/:slug via MCP SDK web-standard transport * feat(agent-mcp-interface): permission gate + navigate tool through MCP * feat(agent-mcp-interface): add read, click, type, attach, submit tools * test(agent-mcp-interface): pin delete-agent-slug-404s-immediately invariant * fix(agent-mcp-interface): plug stub leak + reject non-http navigate + attach traversal * feat(agent-mcp-interface): live integration polish + agent-mcp-manager wiring (#1234) * fix(agent-mcp-ui): clone-from card uses real profiles and hydrates every field * fix(agent-mcp-ui): hide logins step until vault import lands * fix(agent-mcp-ui): pin new-agent rail CTA to viewport bottom * feat(agent-mcp-interface): wire agent-mcp-manager into create + delete * feat(agent-mcp-ui): surface real harness install outcome on the success card * feat(agent-mcp-ui): cover all agent-mcp-manager harnesses + shared HarnessIcon * feat(agent-mcp-ui): real brand marks via @svgl shadcn registry + drop BrowserOS pill * fix(agent-mcp-interface): reconcile harness link on update + regenerate * fix(agent-mcp-ui): invalidate profile caches after create/update/delete/regenerate * fix: handle clone-fetch failure + remove-before-uninstall on delete * feat(agent-mcp-interface): adopt real browser tools + mount inside apps/server (#1235) * feat(server): export browser tool surface + session for cockpit reuse * feat(agent-mcp-interface): adopt @browseros/server real tool catalogue with permission wrapper * feat(server): mount cockpit inside apps/server runtime with mcpUrl migration * test(agent-mcp-interface): pin migrateMcpUrls rewrite + re-install behavior * fix(server,eval): cast around workspace zod version cross-pollination * fix(cockpit): isolate uninstall + catch migration; reject non-http navigate; log run dispatch - migrate-mcp-urls: wrap uninstallForAgent in its own try/catch so a throw there does not skip installForAgent and leave the harness pointing at a dead URL while the profile JSON carries the new one. - cockpit.ts: add .catch on migrateMcpUrls so a top-level rejection (e.g. listFiles hitting EACCES) is logged instead of swallowed as an unhandled promise rejection. - mcp/register: reject javascript:, file:, and data: URLs at the navigate wrapper before the permission gate, restoring the defense-in-depth the old per-tool wrapper had. The real navigate tool's schema is z.string().optional() with no scheme check. - mcp/register: log a warning when the run tool dispatches. A dedicated catalog verb for arbitrary script execution is the proper fix; the log keeps dispatches auditable until that lands. - Integration test: lock the navigate scheme guard with explicit cases for javascript:, file:, and data:. * feat(cockpit-ui): confirm before rotating MCP URL; drop redundant Add-to-harness button The MCP page had two paper cuts: 1. Regenerate URL fired straight on click. Rotating destroys the previously-issued URL and re-installs the harness entry under a new slug, so anywhere the old URL was pasted by hand stops working. Now the button opens a shadcn AlertDialog explaining the impact (auto-reinstall via reconcileHarnessLink, but external paste-ins go dead) and only fires the mutation on confirm. Matches the pattern used by DeleteAgentDialog. 2. The "Add to <harness>" button only flipped a local "Added" badge for 1.8s; it never triggered an install because the install already ran when the agent was created. Removed the button and updated the header + empty-state copy to say so explicitly. * fix(cockpit): mount standalone server under /cockpit prefix The merged-runtime refactor switched the UI client and the harness install URLs to a single shape: `http://127.0.0.1:<port>/cockpit/...`. That works against `createCockpitRoutes` because apps/server mounts the cockpit under `.route('/cockpit', ...)`. The standalone entry point in `src/main.ts` was still serving at the root, so every UI request returned 404. Wraps the Hono `server` in a parent that mounts it under COCKPIT_MOUNT_PREFIX and updates `localServerUrl` to include the prefix. The buildMcpUrl helper now produces the same shape in both runtimes, so harness configs stay valid across a switch. Also runs `migrateMcpUrls` at standalone boot — same sweep the production factory does — so profiles created before this change get their stored mcpUrl + harness install entries rewritten to the new shape on first start. * fix(cockpit-ui): reconcile agent-profiles list after regenerate URL GET /agents is sorted by `updatedAt` DESC server-side, and regenerate bumps that field — so the rotated row needs to jump to the top of the directory and the MCP page. The previous handler only patched `mcpUrl` in place, leaving the sort order stale until the next page load. Swaps the detail-cache invalidation (the GET /agents/:id wire shape doesn't carry slug or mcpUrl, so the invalidation was a no-op) for a list-cache invalidation, while keeping the optimistic `mcpUrl` patch so the new URL still appears without a network round-trip. * feat(agent-mcp-interface): attach to browseros browser over CDP at boot (#1248) * feat(agent-mcp-interface): attach to browseros browser over cdp at boot The standalone cockpit ran the route surface but never set the process-wide BrowserSession, so every MCP tools/call short-circuited with "browser session not connected". Production worked because the merged runtime in @browseros/server called createCockpitRoutes with its live session; standalone had no such hand-off. Mirrors @browseros/server's bootstrap directly: connect a CdpBackend to the configured port, wrap in Browser, hand the session to setBrowserSession at boot. Configurable via the BROWSEROS_COCKPIT_CDP_PORT env var, defaults to 49337 (IANA dynamic / private range, no known collision with registered services). Soft-fails when the browser is not reachable. The cockpit still serves the UI, profile CRUD, harness installs, and tools/list; only tools/call keeps the existing "session not connected" wire shape until the user restarts the cockpit with the browser up. exitOnReconnectFailure: false on the CdpBackend so a transient drop degrades the session instead of killing the cockpit process. CdpClient is injected through BrowserBootstrapDeps so the unit test covers the connect-success, connect-fail, and disconnect-swallows-errors paths without opening a socket. * fix(cockpit): guard signal handler + harden DI shape Three small cleanups from the PR review: - Add an `exiting` guard around the SIGINT/SIGTERM cleanup so a back-to-back delivery (supervisor sends both) does not restart `disconnect()` on an already-closing CDP connection. - Add a `setTimeout(() => process.exit(1), 5000).unref()` kill switch before `disconnect()` so a hung inner `cdp.disconnect()` (half-open socket, network stall) cannot leave the process unkillable except via SIGKILL. - Replace `BrowserBootstrapDeps`'s two independent override fields with a single bundled `inject` object so callers cannot mix a stub `cdpFactory` with the default `buildSession`. The default `buildSession` casts to a real `CdpBackend`, so a partial override would compile but blow up at the first `Browser` call. Tests updated accordingly. * feat(agent-mcp-ui): single source for the MCP endpoint URL shown in agent create/edit (#1328) The wizard, the edit flow, and the MCP directory each constructed the per-agent MCP URL inline. The wizard's pre-save preview returned http://127.0.0.1:9000/mcp/<slug>, which is wrong on both axes: port 9000 is the CDP socket, and the cockpit lives under /cockpit while it borrows apps/server's BrowserSession. Introduce modules/api/mcp-endpoint as the canonical source. The copy widget on /agents/new and /agents/:id/edit, the McpRow in the directory, and the slug parser all flow through it. A TODO at the top of the module marks the temporary cockpit-mount target and lists the three commits the future unmount will need. * feat(cockpit): tab activity registry + homepage feedback loop (PR 1/3) (#1331) * feat(agent-mcp-interface): tab activity registry + GET /tabs/activity route Wraps the existing executeTool dispatch in mcp/register.ts so every successful browser-tool call is recorded against the calling agent and the targeted CDP target id. Failed dispatches and tools without a page arg (tab_groups, windows, run) are skipped. Records live in an in-memory map keyed by targetId; status is derived at read time (active for 5s after the last tool, idle afterwards). Closed tabs are evicted lazily when the next snapshot read finds the pageId no longer maps to the original targetId (pageIds are reused after close). The GET /tabs/activity route surfaces the current snapshot and is mounted into the AppType chain so the UI hono-rpc client picks it up automatically. No server-package edits anywhere; the cockpit reads PageManager via the shared BrowserSession singleton it already owns. * feat(agent-mcp-ui): drive cockpit homepage from real tab-activity polling useTabsActivity polls GET /cockpit/tabs/activity every 1500ms via the existing hono-rpc client; cockpit.data composes that with the existing mocked approvals/handoffs so the screen calls one hook only. Active records become RunningGrid cards (status=running); idle records become RecentActivity rows (status=done) with a relative-time string. Helper file derives a stable per-slug color, parses the site from the URL, and formats the relative timestamp. Mocked useAgents / useRecentActivity stay in place for any other surface that imports them; the homepage just stops consuming them. * fix(cockpit): test-isolation clear() + honest isPending + harness TODO - TabActivityRegistry gains a clear() escape hatch next to size(); the routes/tabs test now calls it in afterEach so a stale record from one test cannot surface in another that re-attaches a session. - useCockpitData isPending now OR-combines tabs/approvals/handoffs so any future caller wiring a spinner sees the actual loading state. - tabsToAgentRows annotates the hardcoded harness with a TODO pointing at the PR-3 profile join so the simplification stays visible. * fix(eval): restore any-typing on ai-sdk onStepFinish params for mixed-zod workspace CI typecheck on the merge commit failed in the eval package: the onStepFinish destructure inherited from main resolves to implicit-any under this branch's workspace where the cockpit pins zod v4 and the server pins zod v3, so the ai-sdk generate() option type widens. Re- introduce the explicit `any` annotations + biome-ignore comments that the branch carried before the merge took main's cleaner shape; the runtime contract is unchanged. Single-agent.ts: `onStepFinish: async (step: any)` with biome-ignore. Tool-loop-executor-backend.ts: typed destructure literal with two biome-ignores on the toolCalls and toolResults any fields. --------- Co-authored-by: shivammittal274 <56757235+shivammittal274@users.noreply.github.com>
Yesterday's gate asked whether a boundary file had changed at all since its issue was filed. That removed the worst of the false "already done" verdicts, and it is the weaker question: browseros-ai#1169 passed it because its file changed for unrelated reasons while the criterion it was actually missing stayed missing. The measurement that matches the claim is when the NAMES arrived. If every identifier the criteria mention was already in the boundary files at the commit that was HEAD when the issue was written, their presence now says nothing about the issue - it says the defect had a name. If one of them arrived afterwards, that is what done looks like. Checked against the cases before landing it, because a stricter gate that re-dispatches finished work is worse than the problem: browseros-ai#1173, browseros-ai#1175 and browseros-ai#1176 name identifiers that were NOT in their boundary file when they were filed, so they stay skipped, correctly. Their work is done. Caught an inversion in my own first draft: on a failed git call it answered "the names predate the issue", which LEFT the already-done verdict standing - the opposite of what its own comment claimed. The decision now lives in QueenEvidencePolicy, where nil means the question was not answered and an unanswered question may not dismiss work. Nine checks pin both directions, the unmeasured case, and the empty ones. The superseded helper is deleted rather than left behind. Gate: REAL_EXIT=0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Asked why there is no working Queen, I asked her. Her own log, one minute old:
queen.choose: Nothing to choose from 13 candidate(s): 5 settled or
waiting on you, 4 looks already done, 2 held by another task,
1 attempts exhausted, 1 no boundary
She is not broken. She ticks, evaluates thirteen candidates, and is allowed to
take none of them. Six merged and ten accepted stand behind her; she has
nothing left she may pick up.
Four of those thirteen are this defect. browseros-ai#1173, browseros-ai#1174, browseros-ai#1175 and browseros-ai#1176 were all
refused as "looks already done" because `handleWorkerFinished`,
`chooseNextOpenIssue`, `autoAcceptIfUnambiguous` and `requestReviewerVerdicts`
were found in rings/SR-00/QueenLocalisation.swift.
The test was `fileContents.contains(identifier)`. Substring containment is not
evidence of anything. Measured against that real file just now:
handleWorkerFinished contains=true declares=false
chooseNextOpenIssue contains=true declares=false
autoAcceptIfUnambiguous contains=true declares=false
requestReviewerVerdicts contains=true declares=false
None of the four is declared there. They are declared in ChatViewModel.swift,
and they appear in the boundary file only inside its narrative header and its
measurement table - which lists them as the INPUTS the narrowing logic is
tested against. So the heuristic read a file's documentation of what it is
tested on, concluded the work was finished, and starved the Queen of a third
of her queue.
`QueenEvidencePolicy.declaresIdentifier` now requires a declaration keyword -
func, var, let, case, enum, struct, class, protocol, typealias, actor,
extension - with the identifier boundary checked on both sides, so
`func handleWorkerFinishedLater` does not answer for `handleWorkerFinished`.
Comment lines are skipped before the search, for the reason the keychain gate
skips them: a line beginning `//` is prose, and prose cannot declare a
function.
The suite goes 21 checks to 35, including the four live names against the
shape of the real file, both boundary directions, and the prefix case.
This does not touch the other nine: five wait on the owner by contract, two
are boundary-held, one has exhausted its attempts, one has no boundary. Those
are correct refusals.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…l anywhere
The dispatch chain had never once executed - it had code review and nothing
else. It has now, three times, in parallel, and nothing secret was installed.
HOW, WITHOUT A KEY. `openai-compatible` is the one provider whose factory
requires a baseUrl and no apiKey. And this repository already proves worker
behaviour by replaying a recorded stream instead of calling a model - that is
what TRIOS_REPLAY_CASSETTE does on the Mac. So a route inside the container
speaks OpenAI chat-completions and streams a scripted reply, and the worker
provider points at it over loopback.
Guarded like everything else: the in-container client presents THIS SERVER'S OWN
token as its apiKey, which openai-compatible puts in the Authorization header.
The same door, used by the process itself, not a new one.
Off unless TRIOS_QUEEN_REHEARSAL is set, and never the silent fallback where a
real key exists. A hive that quietly rehearses instead of working is worse than
one that stops, because it reports success.
WHAT RAN:
16:13:06 Queen dispatch issue=1244 branch="queen-1244" started=true
cut from feat/queen-supervisor
16:16:32 Queen rehearsal turn model="rehearsal"
16:16:32 Queen worker turn finished conversationId=f20e33b7-... issue=1244
Lease taken, stalled dispatches reaped, board read from Postgres, 40 candidates
fetched, queend chose browseros-ai#1244 by its own declared boundary, worktree cut on its
own branch, turn opened, stream consumed, dispatch recorded finished.
SAY PLAINLY WHAT IT IS NOT: no bee thought. The reply is scripted; nothing read
the issue or wrote code. What is proven is the plumbing - which is the part that
had no evidence.
TWO DEFECTS FOUND BY MAKING IT RUN.
`userWorkingDir`, not `workingDirectory`. The schema names it the first way and
ignores unknown keys, so the wrong name was accepted in silence and every bee
would have worked in the shared checkout while its branch lived in a worktree -
edits and branch in different trees, the exact failure a worktree prevents.
git ran as root. The image splits uids and the entrypoint says "git runs as bee;
root does not enter the checkout". Mine did:
fatal: detected dubious ownership in repository at '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/workspace/BrowserOS'
The tempting fix is safe.directory and it is the wrong one: it tells git to stop
minding precisely what the uid split enforces. Dropping to bee through the same
helper every agent shell command uses keeps the split and satisfies git for the
real reason.
PARALLEL BEES UNDER THE QUEEN:
browseros-ai#1240 -> started rings/SR-02/ChatViewModel.swift, free
browseros-ai#1216 -> started docs/queen-choice.md, free
browseros-ai#1176 -> REFUSED QueenLocalisation.swift held by trios#1174
browseros-ai#1286 -> REFUSED a task already exists for it (cancelled)
Two started, two refused for different, issue-specific, correct reasons. Three
worktrees on three branches, all owned by bee.
HOW MANY IN PARALLEL - three numbers, and only the smallest matters:
container capacity dozens (96 concurrent, no degradation; 200/200 in 3.1s)
the Queen's policy 4 (QueenDelegationPolicy.maximumConcurrentWorkers)
what actually binds boundaries - browseros-ai#1174's parked review holds
QueenLocalisation.swift against four issues at once
The hardware was never the limit. It is the boundary ledger and a review with no
verdict.
Still absent, and still not mine to install: a real provider key. The rehearsal
proves the plumbing, not the thinking.
…laptop
ChatRequestBuilder assembles a `messages` array - system prompt, history,
current turn - and posts it under the key `messages`. The server has no such
field. `grep -c messages` on api/types.ts returns 0, ChatRequestSchema is a
plain z.object with no .passthrough(), so zod strips it before any handler runs.
The server accepts, and always has, exactly the two fields being discarded:
types.ts:42 userSystemPrompt -> prompt.ts:649
types.ts:82 previousConversation -> chat-service.ts:378
So on every send: the Queen's prompt, every bee's prompt, the reviewer's prompt
and the whole conversation history were computed and thrown away. All three ran
as the generic assistant with an empty preferences block - which explains why a
bee behaves like a stranger to this project however carefully its brief is
written. The part that told it what it was never arrived.
The codebase had already recorded the cost without the cause: a comment at
ChatViewModel.swift:12320 describes a bee that "found an unrelated old checkout
under ~/gitbutler and edited that instead, so its branch here stayed empty". The
sentence that prevents exactly that was in the prompt that never travelled.
THE CLOUD BEE WAS WORSE OFF. Three things, each one variable away in the same
function: its boundary (computed by queend, used to refuse three other
candidates, never told to the bee it constrains); the issue body (the brief
opened with "Read the issue first" - impossible: no gh in the image, and
GITHUB_TOKEN deliberately excluded from the shell allowlist, so the only
description it got was the number); and any identity at all. All three now
travel with the dispatch.
Also here:
- the provider key travels with the choice. Without it a deployment WITH a key
cleared the credential precheck and died at the chat route saying the key
was absent.
- queend refused the whole question once anything was in flight: Postgres
hands back milliseconds and Swift's .iso8601 will not take a fraction, then
acceptanceCriteria was missing. The record is now built from the type's
field list rather than from what seemed interesting.
- z.ai default model glm-4.6 -> glm-5.3, confirmed live: three bees running.
- /queen/dashboard, in t27.ai's palette on a phi scale. The shell is unguarded
and stateless; the data stays behind the token. A token in the URL was the
obvious alternative and is the one thing that must not happen.
- the Dockerfile said a volume mounts at /workspace. None was ever created -
df says overlay, mount says nothing, reflog says "clone:", and railway
volume list has one volume, on Redis. Every redeploy re-clones, and it has
already destroyed e52f41ad, a real commit by a real model. The comment now
says what is true.
NOT DONE: `make queen-core-sync` is red on QueenLocalisation.swift, which
another agent is editing uncommitted in this shared tree (browseros-ai#1176 work). Syncing
the Linux copy would land their half-finished change with mine, so it stays red
and stays theirs.
Gates: make-dollars, recipe-backticks, empty-sources, queen-core (14 files,
3950 lines), sources-drift (201), warnings 0/196, 268 server tests.
…he page
"Не открывается" was fair, and so was "почему не проверяешь результат". I
verified the container and never once opened the page on the machine the data
lives on. Three defects, none of which a status code would have shown:
ONE. The tree looked for its JSON at `${cwd}/../../.trinity/...`. The server
runs from `trios/agent-server`, so the project root is ONE level up, not two -
it reached `BrowserOS/.trinity`, which does not exist. The page would have
rendered "no tree found" on the very machine the file was written on.
TWO. The dashboard served literal `—` where an em dash belonged - visible
in the title, in the token hint, in the footer. The source carried a raw non-ASCII
character (an L3 violation in its own right) and the pipeline escaped it on the
way out. Every non-ASCII character in both pages is now an HTML entity, which is
both the fix and what L3 asked for. biome had flagged the String.raw that made
it visible and I had ignored the warning.
THREE. The board rendered six empty columns against a database with no swarm in
it, which looks exactly like a quiet swarm. It now says which of the two it is -
the same distinction the tree draws between a missing file and an empty one.
THE BOARD ITSELF. Columns are the Queen's OWN states, not invented ones: a board
whose vocabulary the logs do not use teaches its reader the wrong words.
backlog 25 | blocked 2 | running 2 | in review 1 | done 10 | dropped 17
57 cards, every one clicking through to its GitHub issue, and it tells the
session's story at a glance: browseros-ai#1174 sits in review and holds
rings/SR-00/QueenLocalisation.swift against browseros-ai#1176 and browseros-ai#1175, both shown as
"held by browseros-ai#1174". That is the thing that has been quietly binding this swarm,
and it took a column layout to make it obvious.
`queen_issues` caches what GitHub showed the tick, so drawing the board costs
none of the anonymous 60/hour budget. The boundary is parsed when the issue is
stored, by the same two-heading rule the Queen uses - a second parser in the
page would have agreed until one of them was edited.
Verified by looking: six columns, no horizontal overflow, 57 cards, every link
absolute to github.com. And a fourth defect caught that way - the RUNNING
column, the one an operator reads first, showed bare numbers because a dispatch
the app never saw has no registry task to borrow a title from. It now borrows
from the issue list.
…ased the issue "Почему running 0" has an innocent answer and the digging found a real one. RUNNING 0 IS NOT A FAULT. The tick fires every 30 minutes, a turn takes about ten, so most of any half hour has nothing running. The last round was 16 minutes before the question. THE REAL FAULT was on the bee's own branch: a994a8b 12 minutes ago sixth verification record for browseros-ai#1244 - all checks hold 137dad0 38 minutes ago fifth verification record for browseros-ai#1244 - all checks hold 5b7f473 65 minutes ago fourth verification record for browseros-ai#1244 - all checks hold 948449b 2 hours ago third verification record for browseros-ai#1244 - all checks hold The work was finished hours ago. Every thirty minutes the Queen chose browseros-ai#1244 again, a bee arrived, found the job already done, verified it, and committed a record saying so. Six times, six model turns, on one issue. Because finishing RELEASED it. The in-flight query was `finished_at IS NULL`, so the moment a turn ended the issue was choosable again - and it was still open on GitHub, because nothing lands a bee's branch. A reaped dispatch must release its issue (its container died, nothing was finished). A finished one must NOT: it holds until somebody judges it, exactly as awaitingReview does on the Mac. The board shows it in review now rather than letting it vanish and reappear. Verified after the change: browseros-ai#1244 moved out of backlog into review. Round seven will not take it. THE BLOCKERS, measured rather than guessed: 24 of 27 backlog issues declare no boundary, so the Queen refuses them - she cannot reserve files for a task that has not said which files it touches. Her real choosable pool is THREE, not 27. browseros-ai#1174 sits in review holding rings/SR-00/QueenLocalisation.swift against browseros-ai#1176 and browseros-ai#1175. awaitingReview is not terminal. Only a verdict frees it. MAGAZINE HEADLINES, as asked: a display scale (clamp 2.6-5.2rem, -0.045em tracking, 0.95 leading), a letterspaced kicker above and a hairline rule under, on all four pages. AND THE FORM WOULD NOT HIDE. `el.hidden` was true and the box stayed on screen: the hidden attribute is a UA rule and loses to any author display rule, so .auth{display:flex} beat it. Measured - hidden=true, computed display=flex, 57 cards behind it. Every page now carries [hidden]{display:none !important}. FOURTH BACKTICK OF THE DAY, in the CSS comment explaining that fix, inside a template literal - 16 typecheck errors across four files. The SQL-only gate let it through, so it now covers page shells too. Its first version then reported three closing backticks as offences, because a shell closes at the end of the last markup line rather than on a line of its own; that shape is pinned by a test so the noise cannot come back. typecheck 0, 280 tests across 27 files.
…s nothing The first thing the new headline did was contradict itself, on the live board, within a minute of shipping: Nothing is running, and 1 issue is ready. Her own reason, from the last round: nothing to choose Both sentences came off the same page. The board's list of who is HOLDING a file was built from the app's registry alone, so browseros-ai#1175 looked free - while the round had refused it with rings/SR-00/QueenLocalisation.swift held by gHashTag/trios#1176 and browseros-ai#1176 is a CLOUD dispatch, which the registry knows nothing about. The page and the Queen were reading different boards, and the page's number was the one on the front in the largest type. Holders now come from both sides, shaped as tasks so the same `stillHoldsBoundary` clock applies to each: a running dispatch holds, a finished one holds as `awaitingReview` and ages out at 48 hours, exactly as the policy says. Proven by reverting the holder list to registry-only and watching 'puts the issue in blocked, not backlog' go red. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ros-ai#1176) (#307) Second pass. The first pass put the self-check in and then watched it change nothing - and said so - because the check compared a range's first line against the name extracted from that same line. Tautology: rule 1's window is clamped to the declaration line by construction, so the guard could never fire and removing it could not break anything. Criterion 4 was unmeetable as wired. The re-measurement found what actually produces wrong ranges on today's file, and it is not rule 1: 1. maskComments opened block comments inside string literals. A branch glob quoted at line 6864 - "No empty queen/* branch ..." - blanked the literal view from there to the end of the file: 6 726 lines of evidence invisible. Corroboration saw no mentions past the glob, browseros-ai#1158's guard and neighbour tied at zero, rule 1 went silent, and rule 4 answered with 6005-6304 - the middle of handleWorkerFinished, a function that spec never named. The confident wrong range browseros-ai#1176 is about, produced by the wrong rule through a blind view. maskComments tracks strings now, using the same model maskCommentsAndStrings always had. 2. The self-check moved to where ranges are born: every rule's answer passes through one finishing step - the first line of the returned range must declare a function, and the name rule must land on the very name it matched. A window that capping slid into the middle of a body begins nowhere in particular, and beginning nowhere in particular is silence. The check is live: queen.review.verdicts sits 322 lines into the 376-line requestReviewerVerdicts, its window would start at 7508, and the check silences it. Delete the two guards in finished() and that replay row goes red - verified from both sides; that is the criterion-4 property. The re-recording (identifiers unchanged; the boundary file moved to 13 590 lines): browseros-ai#1158 points at autoAcceptIfUnambiguous - the neighbour corroborates 0, the subject 2, measured not assumed; browseros-ai#1117 points at requestReviewerVerdicts (7432-7731, beginning at its declaration); browseros-ai#1156's clue moved upstream into settleCharacterCountVerdicts and the recording follows the clue, not the old address; browseros-ai#1166 unchanged at chooseNextOpenIssue. Expected gains declarationOrSilence - the issue's own acceptance, verbatim: point at the named function, or say nothing; a range that begins anywhere else is a FAIL. A returned range now counts as ok only when it BEGINS at the named function's declaration line and stays inside its extent - "somewhere inside" was the loophole the mid-body ranges walked through. One follow-up outside this task's boundary: the Linux module copy at agent-server/queen-core/Sources/QueenCore/QueenLocalisation.swift is byte-compared by make queen-core-sync and now differs until someone copies rings/SR-00 over it. The boundary names one file, so this task did not touch it. Co-authored-by: Trinity Bee <bee@trinity.local>
Summary
packages/browseros-agentfallow report from 17 issues to 4 (bun fallowon the prior main was failing because of dead-code findings introduced by feat: one-click install of BrowserOS as MCP into installed agents + URL auto-sync #1168).mcp-managerpublic surface, narrows two internal helpers, exports one props interface to fix a real private-type leak, removes two genuinely unused devDeps (dotenv,picocolors).RemoteHermesVmStatus,PostTurnResult,RemoteHermesBootPillProps,SocketState) are all in the brand-newremote-hermes/*surface from feat(agent): Remote Hermes provider — Cloudflare-managed Fly VM runtime #1174 and intentionally out of scope here. They can be cleaned up in a follow-up by whoever owns that surface.What changed
apps/server/src/lib/mcp-manager/index.ts— slim the barrel to only what's actually imported through it:humaniseInstallError,installInto,listAgents,uninstallFrom,reconcileUrl,resetMcpManagerForTesting,setMcpManagerForTestingBROWSEROS_MCP_SERVER_NAME,BROWSEROS_MCP_STDIO_SERVER_NAME,getMcpManager,ReconcileUrlInput,InstallAgentResult,McpAgentId,McpAgentRow,ReconcileResult,UninstallAgentResult(all of these are only used by sibling files insidemcp-manager/itself, which import directly from./manager,./reconcile,./types— never via the barrel)apps/server/src/lib/mcp-manager/service.ts:BROWSEROS_SERVER_NAMESandplanFornarrowed to module-private (both are only used insideservice.ts)McpAgentIdentifiertype alias deleted (zero consumers)McpAgentIdimport that the alias kept aliveapps/agent/screens/mcp-settings/IntegrationsSection.tsx:IntegrationsSectionPropsis nowexport interfaceso the exportedIntegrationsSectioncomponent no longer references a private type (purely additive — no consumer change)apps/agent/screens/mcp-settings/integrations-section.helpers.ts:AGENT_PRESENTATIONnarrowed to module-private (only used bypresentationFor()in the same file)package.json:dotenvandpicocolorsfromdevDependencies. Neither is imported anywhere in the package; the agent README explicitly notesdotenvisn't needed since env files are loaded automatically.bun.lockregeneratedRisk
packages/browseros-agenttree to confirm zero external consumers (including the agent UI side and the test suites).exportfrom a const/function that's only used in the same file is purely a visibility narrowing.exportto an interface is purely additive.Test plan
bun fallowinpackages/browseros-agent— was 17 issues, now 4 (all four are inremote-hermes/*, out of scope)bun run typecheck— clean (@browseros/agent,@browseros/server,@browseros/eval,@browseros/cdp-protocol,@browseros/build-toolsall pass)bun test apps/server/tests/lib/mcp-manager/— 15/15 pass (bothreconcile.test.tsandservice.test.ts)