Skip to content

fix: bootstrap db schema when migrations are missing - #1244

Merged
Nikhil (shadowfax92) merged 5 commits into
mainfrom
fix/0617-1804-wombat
Jun 18, 2026
Merged

Nikhil (shadowfax92) merged 5 commits into
mainfrom
fix/0617-1804-wombat

Conversation

@shadowfax92

Copy link
Copy Markdown
Contributor

Summary

  • add a temporary SQLite schema bootstrap when Drizzle migration files are missing from packaged resources
  • validate migration folders before passing them to Drizzle so empty/partial directories fall back cleanly
  • seed the current Drizzle migration history and assert exact hashes/timestamps in DB initialization tests

Verification

  • bun test apps/server/tests/lib/db/index.test.ts apps/server/tests/lib/agents/db-agent-store.test.ts apps/server/tests/lib/clients/oauth/token-store.test.ts
  • bun run --filter @browseros/server typecheck
  • bun run lint
  • git diff --check
  • bun --env-file=.env.development test --preload=./tests/helpers/test-env.ts ./tests/tools/bookmarks.test.ts ./tests/tools/filesystem/bash.test.ts

Note: a full bun run check hit a transient CDP/test-harness failure in the tools group; after cleanup, the failed bookmark/filesystem bash files passed cleanly.

@shadowfax92

Copy link
Copy Markdown
Contributor Author

Greptile (@greptileai) review

@shadowfax92

Copy link
Copy Markdown
Contributor Author

Claude (@claude) review

@github-actions github-actions Bot added the fix label Jun 18, 2026
@claude

claude Bot commented Jun 18, 2026 •

Copy link
Copy Markdown

Claude finished Nikhil (@shadowfax92)'s task in 3m 18s —— View job


Code Review

  • Gather context and read changed files
  • Review client.ts changes
  • Review index.test.ts changes
  • Post final review feedback

Good overall structure. The transaction wrapping, idempotent DDL (CREATE TABLE IF NOT EXISTS), and the journal validation are all solid. A few things worth addressing before merging:


client.ts

1. Silent fallback on explicit migrationsDir — behavior regression ⚠️

client.ts:69-73

if (options.migrationsDir) {
  if (hasCompleteMigrationSet(options.migrationsDir)) {
    return options.migrationsDir
  }
  return null   // ← was a thrown error before
}

The old code threw if an explicit path was invalid. Now it silently falls back to bootstrap. If a developer typos the path (or the build step puts migrations in the wrong place), they'll get a silently bootstrapped DB with no signal. At minimum, a logger.warn(...) here would make production failures visible. Without it, this change erases an intentional guard. Fix this →

2. No log on bootstrap path ⚠️

client.ts:52-54

} else {
  bootstrapCurrentSchema(sqlite)
}

When this branch is taken in production there's no trace in the logs. If a user reports a startup issue, there's no way to know the DB was bootstrapped rather than migrated. A single logger.info('Drizzle migrations unavailable; bootstrapping current schema.') here would go a long way. Fix this →

3. SERIAL PRIMARY KEY is not SQLite syntax ⚠️

client.ts:264-270

CREATE TABLE IF NOT EXISTS __drizzle_migrations (
  id SERIAL PRIMARY KEY,   -- ← PostgreSQL type
  hash text NOT NULL,
  created_at numeric
)

SQLite has no SERIAL type — it assigns NUMERIC affinity to unknown type names. This works (tests pass), but INTEGER PRIMARY KEY would be the correct SQLite idiom for an autoincrement rowid alias. More importantly, when migrate() is called against a bootstrapped DB on second run, Drizzle's internal migrate function reads and writes this table. If Drizzle uses INTEGER PRIMARY KEY in its own DDL for this table, the schemas will differ between the bootstrap and migration paths. Worth cross-checking against drizzle-orm/bun-sqlite's migrate source. Fix this →

4. String interpolation in SQL (low risk, style nit)

client.ts:149-156

sqlite.exec(`
  INSERT INTO __drizzle_migrations ("hash", "created_at")
  SELECT '${migration.hash}', ${migration.createdAt}
  ...
`)

currentMigrationHistory is a hardcoded constant so there's no injection risk here. Still, Bun's sqlite.prepare(...).run(...) API would be cleaner and safer by convention, especially since this DDL table may be extended in future.

5. Maintenance hazard: currentSchemaStatements and currentMigrationHistory drift

This is called out in the TODO, so it's acknowledged — but there's no mechanism preventing them from going stale when a new migration is added. The hasCompleteMigrationSet check guards against using an incomplete migrations dir, but nothing guards against forgetting to update the bootstrap constants. A comment in drizzle.config.ts or a test that asserts currentMigrationHistory matches the actual journal entries would catch this. Fix this →


index.test.ts

Overall: test coverage is thorough and well-structured. The expectCurrentSchema helper is clean. The four new scenarios cover the important paths.

6. Weak assertion in skips empty packaged migration resources

index.test.ts:82-83

expect(handle.migrationsDir).not.toBe(packagedMigrationsDir)

This passes whether handle.migrationsDir is null (bootstrap) or sourceMigrationsDir (source tree found). Since the intent of this test is to verify the empty packaged dir is skipped in favor of fallback, a more explicit assertion would be:

// If source migrations are available in the test environment:
expect(handle.migrationsDir).not.toBe(packagedMigrationsDir)
// Or if you want to assert bootstrap was used:
expect(handle.migrationsDir).toBe(null)

The current assertion does verify the empty packaged dir is not used, but the intent could be clearer.

7. Removed error test without replacement coverage

The old 'fails clearly when an explicit migration directory is missing' test was removed. The new 'bootstraps the current schema when migration files are unavailable' test covers the happy path, but there's no test asserting that a warning/log is emitted (which ties back to finding #1 above). Once logging is added, a test verifying that the warning fires on explicit-path fallback would be valuable.


Summary

The core approach is sound and solves a real packaging gap. The three issues worth addressing before merge are:

  1. Log (or warn) when bootstrap is used, especially on the explicit-migrationsDir fallback path
  2. Verify SERIAL vs INTEGER PRIMARY KEY against Drizzle's internal schema
  3. Add a drift guard for currentMigrationHistory vs the actual journal

@github-actions

github-actions Bot commented Jun 18, 2026 •

Copy link
Copy Markdown
Contributor

✅ Tests passed — 1314/1320

Suite Passed Failed Skipped
✅ agent 184/184 0 0
✅ build 18/18 0 0
✅ eval 91/91 0 0
✅ server-agent 296/296 0 0
✅ server-api 135/135 0 0
✅ server-browser 9/9 0 0
✅ server-integration 10/10 0 0
✅ server-lib 253/254 0 1
✅ server-root 52/55 0 3
✅ server-tools 266/268 0 2

View workflow run

@greptile-apps

greptile-apps Bot commented Jun 18, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds a schema-bootstrap fallback for packaged builds where Drizzle migration files are absent, and introduces hasCompleteMigrationSet to validate a migrations directory before handing it to Drizzle. The DbHandle.migrationsDir type widens to string | null to signal which path was taken.

  • resolveMigrationsDir now returns null instead of throwing when no complete migration set is found, causing openBrowserOsDatabase to call bootstrapCurrentSchema — an inline DDL sequence that creates all tables and seeds the migration-history table.
  • hasCompleteMigrationSet parses _journal.json, verifies every tag in currentMigrationHistory is present, and confirms each .sql file exists — incomplete or empty directories fall back cleanly.
  • Four new tests cover missing-dir bootstrap, empty-dir bootstrap, empty-packaged-resources skip, and idempotent re-open after bootstrap.

Confidence Score: 3/5

Safe to merge once the __drizzle_migrations DDL is corrected; the bootstrap logic itself is sound but the wrong column type creates an ongoing schema mismatch.

The bootstrap path correctly guards against re-running migrations and the new tests cover the main scenarios well. However, the __drizzle_migrations table is created with id SERIAL PRIMARY KEY — SQLite does not treat SERIAL as an auto-incrementing integer, so every row inserted by bootstrap and by any subsequent Drizzle migrator run will have a NULL id. This diverges from the table structure Drizzle generates and expects.

The bootstrapCurrentSchema function in client.ts, specifically the __drizzle_migrations DDL at the bottom of currentSchemaStatements.

Important Files Changed

Filename Overview
packages/browseros-agent/apps/server/src/lib/db/client.ts Adds schema-bootstrap fallback and migration-set validation; __drizzle_migrations DDL uses SERIAL PRIMARY KEY which is not an auto-incrementing type in SQLite (should be INTEGER PRIMARY KEY AUTOINCREMENT NOT NULL), creating a schema mismatch with what Drizzle itself generates.
packages/browseros-agent/apps/server/tests/lib/db/index.test.ts Replaces the "throws on missing dir" test with four new scenarios covering bootstrap, empty-dir, empty-packaged-resources, and idempotent re-open; coverage is thorough and the expected migration hashes match the constants in client.ts.

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart TD
    A[openBrowserOsDatabase] --> B[resolveMigrationsDir]
    B --> C{options.migrationsDir\nprovided?}
    C -- Yes --> D{hasCompleteMigrationSet?}
    D -- Yes --> E[return migrationsDir]
    D -- No --> F[return null]
    C -- No --> G[iterate candidates:\nresourcesDir / sourceMigrationsDir]
    G --> H{hasCompleteMigrationSet?}
    H -- Yes --> E
    H -- No --> F
    E --> I{runMigrations !== false?}
    F --> I
    I -- Yes + dir not null --> J[migrate via Drizzle]
    I -- Yes + dir null --> K[bootstrapCurrentSchema]
    K --> L[CREATE TABLE IF NOT EXISTS\nfor each table + index]
    L --> M[INSERT migration hashes\nWHERE NOT EXISTS]
    J --> N[DbHandle returned]
    K --> N
    I -- false --> N
Loading
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
    A[openBrowserOsDatabase] --> B[resolveMigrationsDir]
    B --> C{options.migrationsDir\nprovided?}
    C -- Yes --> D{hasCompleteMigrationSet?}
    D -- Yes --> E[return migrationsDir]
    D -- No --> F[return null]
    C -- No --> G[iterate candidates:\nresourcesDir / sourceMigrationsDir]
    G --> H{hasCompleteMigrationSet?}
    H -- Yes --> E
    H -- No --> F
    E --> I{runMigrations !== false?}
    F --> I
    I -- Yes + dir not null --> J[migrate via Drizzle]
    I -- Yes + dir null --> K[bootstrapCurrentSchema]
    K --> L[CREATE TABLE IF NOT EXISTS\nfor each table + index]
    L --> M[INSERT migration hashes\nWHERE NOT EXISTS]
    J --> N[DbHandle returned]
    K --> N
    I -- false --> N
Loading
Prompt To Fix All With AI
Fix the following 2 code review issues. Work through them one at a time, proposing concise fixes.

---

### Issue 1 of 2
packages/browseros-agent/apps/server/src/lib/db/client.ts:264-268
The `__drizzle_migrations` DDL uses `SERIAL PRIMARY KEY`, but SQLite has no native `SERIAL` type — it carries NUMERIC affinity and is **not** an alias for the rowid. Drizzle creates this table as `id integer primary key autoincrement not null`; deviating means the `id` column will be NULL for every row inserted by both the bootstrap code and any subsequent Drizzle migrator run. SQLite allows multiple NULLs in a non-INTEGER primary key due to a long-standing quirk, so writes won't error today, but if Drizzle ever issues an `ORDER BY id` or expects non-null ids the results will be wrong. Use the exact DDL Drizzle generates to guarantee schema compatibility.

```suggestion
    CREATE TABLE IF NOT EXISTS __drizzle_migrations (
      id INTEGER PRIMARY KEY AUTOINCREMENT NOT NULL,
      hash text NOT NULL,
      created_at numeric
    )
```

### Issue 2 of 2
packages/browseros-agent/apps/server/src/lib/db/client.ts:69-73
**Silent fallback on explicit `migrationsDir`**

When a caller explicitly passes `migrationsDir`, a missing or incomplete directory now silently falls back to schema bootstrap instead of throwing. The previous behaviour gave a clear diagnostic (`Drizzle migrations directory not found`) — the silent swallow makes it easy for a misconfigured test or deployment to bootstrap an empty DB without any warning, then fail later with opaque "table not found" errors rather than at the point of misconfiguration.

Reviews (1): Last reviewed commit: "test: assert db fallback migration histo..." | Re-trigger Greptile

Comment thread packages/browseros-agent/apps/server/src/lib/db/client.ts
Comment thread packages/browseros-agent/apps/server/src/lib/db/client.ts
@greptile-apps

greptile-apps Bot commented Jun 18, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds a bootstrap fallback that creates the SQLite schema directly from hardcoded DDL statements when Drizzle migration files are absent from a packaged build, and updates resolveMigrationsDir to return null (instead of throwing) when the migration directory is missing or incomplete.

  • client.ts: resolveMigrationsDir now returns string | null; a new hasCompleteMigrationSet validator gates migration folder use; bootstrapCurrentSchema creates all tables and seeds __drizzle_migrations rows in a transaction when no valid folder is found.
  • index.test.ts: Replaces the "throws on missing directory" test with four new cases covering missing, empty, and resource-dir migration paths, plus an idempotent re-open check that validates the bootstrap → normal-migration round-trip.

Confidence Score: 3/5

The bootstrap path will silently produce a __drizzle_migrations table whose id column has no auto-increment behaviour, diverging from the schema Drizzle creates normally. Every database that passes through the bootstrap will carry this divergence forward even after packaging is fixed, because Drizzle's later CREATE TABLE IF NOT EXISTS becomes a no-op.

The SERIAL PRIMARY KEY in the __drizzle_migrations DDL is not auto-incrementing in SQLite — only INTEGER PRIMARY KEY triggers that behaviour. All bootstrapped rows get id = NULL, which SQLite's PRIMARY KEY constraint permits (a historical quirk). When the app later transitions to real migrations, Drizzle silently reuses the existing table, so all future migration rows also receive id = NULL. This is a persistent, silent schema divergence on every affected installation.

packages/browseros-agent/apps/server/src/lib/db/client.ts — the __drizzle_migrations DDL and the SQL string-interpolation in bootstrapCurrentSchema both warrant a closer look.

Important Files Changed

Filename Overview
packages/browseros-agent/apps/server/src/lib/db/client.ts Adds bootstrap fallback schema + migration history when Drizzle migration files are absent; resolveMigrationsDir now returns null instead of throwing. The __drizzle_migrations DDL uses SERIAL PRIMARY KEY, which SQLite does not recognise as auto-increment, producing a persistent schema divergence from what Drizzle's migrator creates.
packages/browseros-agent/apps/server/tests/lib/db/index.test.ts Replaces the "throws on missing directory" test with four new scenarios covering bootstrap from missing/empty/resource dirs and idempotent re-open; expectCurrentSchema validates exact tables and migration hashes. Coverage is thorough.

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart TD
    A[openBrowserOsDatabase] --> B[resolveMigrationsDir]
    B --> C{options.migrationsDir set?}
    C -->|Yes| D{hasCompleteMigrationSet?}
    D -->|Yes| E[return migrationsDir]
    D -->|No| F[return null]
    C -->|No| G[Build candidate list\nresourcesDir + sourceMigrationsDir]
    G --> H{Any candidate\npasses hasCompleteMigrationSet?}
    H -->|Yes| E
    H -->|No| F
    E --> I[migrate drizzle db\nfrom folder]
    F --> J[bootstrapCurrentSchema\nsqlite exec BEGIN]
    J --> K[CREATE TABLE IF NOT EXISTS\nfor each schema statement]
    K --> L[INSERT INTO __drizzle_migrations\nfor each migration in history]
    L --> M[COMMIT]
    I --> N[return DbHandle\nmigrationsDir = path]
    M --> O[return DbHandle\nmigrationsDir = null]
Loading
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
    A[openBrowserOsDatabase] --> B[resolveMigrationsDir]
    B --> C{options.migrationsDir set?}
    C -->|Yes| D{hasCompleteMigrationSet?}
    D -->|Yes| E[return migrationsDir]
    D -->|No| F[return null]
    C -->|No| G[Build candidate list\nresourcesDir + sourceMigrationsDir]
    G --> H{Any candidate\npasses hasCompleteMigrationSet?}
    H -->|Yes| E
    H -->|No| F
    E --> I[migrate drizzle db\nfrom folder]
    F --> J[bootstrapCurrentSchema\nsqlite exec BEGIN]
    J --> K[CREATE TABLE IF NOT EXISTS\nfor each schema statement]
    K --> L[INSERT INTO __drizzle_migrations\nfor each migration in history]
    L --> M[COMMIT]
    I --> N[return DbHandle\nmigrationsDir = path]
    M --> O[return DbHandle\nmigrationsDir = null]
Loading
Prompt To Fix All With AI
Fix the following 2 code review issues. Work through them one at a time, proposing concise fixes.

---

### Issue 1 of 2
packages/browseros-agent/apps/server/src/lib/db/client.ts:264-268
**`SERIAL` is not a valid SQLite auto-increment type** — SQLite only treats `INTEGER PRIMARY KEY` as an alias for the rowid (auto-increment). Any other type name, including `SERIAL`, gets numeric affinity but no auto-increment behavior. When this `INSERT` omits `id`, SQLite assigns `NULL` instead of a generated integer. Because SQLite's `PRIMARY KEY` constraint does not imply `NOT NULL` (except for `INTEGER PRIMARY KEY`), multiple rows can each have `id = NULL`, creating an internally inconsistent table. Drizzle's own migrator creates this table as `INTEGER PRIMARY KEY AUTOINCREMENT NOT NULL`; using `SERIAL` here permanently diverges the schema for any database that went through the bootstrap path, since Drizzle's later `CREATE TABLE IF NOT EXISTS` is a no-op once the table already exists.

```suggestion
    CREATE TABLE IF NOT EXISTS __drizzle_migrations (
      id INTEGER PRIMARY KEY AUTOINCREMENT,
      hash text NOT NULL,
      created_at numeric
    )
```

### Issue 2 of 2
packages/browseros-agent/apps/server/src/lib/db/client.ts:148-157
The `hash` value is interpolated directly into the SQL string. Even though these come from the hardcoded `currentMigrationHistory` constant (so there is no injection risk today), `bun:sqlite` supports prepared statements and parameterised queries via `sqlite.prepare()`. Using raw string interpolation for SQL values is fragile if `currentMigrationHistory` is ever extended with values from a less controlled source, and it makes the intent less clear. Prefer `sqlite.prepare()` with bound parameters so the separation between code and data is always explicit.

```suggestion
    const insertMigration = sqlite.prepare(`
      INSERT INTO __drizzle_migrations ("hash", "created_at")
      SELECT ?, ?
      WHERE NOT EXISTS (
        SELECT 1 FROM __drizzle_migrations
        WHERE created_at = ?
      )
    `)
    for (const migration of currentMigrationHistory) {
      insertMigration.run(migration.hash, migration.createdAt, migration.createdAt)
    }
```

Reviews (2): Last reviewed commit: "test: assert db fallback migration histo..." | Re-trigger Greptile

Comment thread packages/browseros-agent/apps/server/src/lib/db/client.ts
@shadowfax92
Nikhil (shadowfax92) enabled auto-merge (squash) June 18, 2026 03:26
@shadowfax92
Nikhil (shadowfax92) merged commit 3e6ae0a into main Jun 18, 2026
22 checks passed
Vasilev Dmitrii (gHashTag) added a commit to gHashTag/BrowserOS that referenced this pull request Aug 29, 2026
… own work

TWO CORRECTIONS, one of them to a claim I made twice.

I reported "0 open issues, confirmed twice, by two independent methods". Both
methods pointed at `gHashTag/BrowserOS` - the monorepo's git remote. The Queen's
issues live in `gHashTag/trios`, which every slug in her own registry says
plainly: `gHashTag/trios#1286`. Two methods against the same wrong repository
are not two confirmations. They are one mistake, checked twice.

The right repository has FORTY open issues.

What gave it away was `issue 1090 -> 404` from inside the container, for an epic
that certainly exists. A 404 for something known to be there is worth more than
a 0 for something assumed absent.

With the repository right, the round runs the whole way:

  browseros-ai#1244  branch=queen-1244  started=False   15:13:02
    no provider credential in this deployment - set one of ZAI_API_KEY, ...
    Only the operator can.

Lease taken, registry read, 40 candidates fetched, and `queend` picked browseros-ai#1244
because its declared boundary - BR-OUTPUT/QueenTabView.swift - is held by
nobody. Chosen by the Queen's own policy, compiled for Linux, against a boundary
parsed by the same rule the app uses.

The credential-first ordering is now measured rather than read:

  /workspace/BrowserOS  338a8c6 [feat/queen-supervisor]
  branches: 0
  .worktrees: 0

A refusal that left no branch and no directory. The round costs nothing when it
cannot proceed.

THE SECOND CORRECTION: the tick could not see its own work.

The registry mirror is written by the app and knows nothing about what the tick
started. A round would have chosen an issue, dispatched a bee, and thirty
minutes later found the same issue unclaimed and dispatched another - forever,
each bee cutting a branch over the last one's. The symptom would have been a
swarm that looks busy beside a registry that never grows.

In-flight dispatches are now shaped as running tasks and merged into the board,
so both guards apply: "a task already exists for it" and the boundary conflict
check. Which is why `queend` now returns `chosenPaths` and the dispatch row
stores them: a task holding no paths holds nothing against anyone.

Left: a provider key on the deployment, and that is the whole of it for a bee to
run in the cloud. The refusal names every variable that would do it.

Gates: queen-core-sync green, tsc clean, queend exercised on all four decisions.
Vasilev Dmitrii (gHashTag) added a commit to gHashTag/BrowserOS that referenced this pull request Aug 29, 2026
I reproduced, one layer down, the exact defect this session named as what
actually binds the swarm.

`dispatchBee` recorded started=true and stopped. Nothing wrote an ending, so the
row stayed in flight for ever and the in-flight query fed it back into every
future round as a running task holding BR-OUTPUT/QueenTabView.swift. That issue
could never be chosen again.

`awaitingReview` is not terminal, so a parked task holds its boundary
permanently; browseros-ai#1286 held one for five days. I read that, wrote it into the
record, and then built it again.

Three fixes:

  - the stream's end writes the outcome, INCLUDING on the error path. A bee
    whose connection dropped is a bee that is not working, and treating it as
    running is how the boundary leaks.
  - a reaper releases dispatches with no completion after two hours, because a
    container redeployed mid-turn takes its streams with it and leaves nobody
    to write the ending. It runs BEFORE the board is read - a board read first
    is a board with phantom work on it, and the round would skip a candidate on
    behalf of a bee that died an hour ago.
  - a refusal is recorded as already ended. "Refused an hour ago" and "running
    for an hour" are the two states an operator most needs to tell apart.

    browseros-ai#1244  started=False  ended=2026-08-29T15:43:46
       no provider credential in this deployment - set one of ZAI_API_KEY, ...

TWO OUTAGES, BOTH MINE, both worth the words.

A backtick in a comment. The migration SQL is a JS template literal, and I put
backticks around a symbol name inside a SQL comment, which ended the literal:

  ReferenceError: awaitingReview is not defined
    at /app/apps/server/src/lib/db/pg-migrate.ts:133:1

The server 502'd on deploy. A note about a stuck boundary took the service down.
The comment now says so, in the string.

A silent no-op, twice. I patch with Python .replace(); twice I omitted the
assertion, the anchor had been reformatted by biome, and the edit did nothing
while its neighbours landed - producing first a drain() called with four
arguments and defined with two, and then a route that SELECTed finished_at and
never put it in the response. The second cost two deploys spent debugging a
column that was correct all along, because a missing key and a null value look
identical from outside. Every anchor gets an assert.

Gates: make-dollars, empty-sources, queen-core (14/3907), queen-core-sync,
sources-drift (201), warnings 0/196, 268 server tests across 26 files.
Vasilev Dmitrii (gHashTag) added a commit to gHashTag/BrowserOS that referenced this pull request Aug 29, 2026
…l anywhere

The dispatch chain had never once executed - it had code review and nothing
else. It has now, three times, in parallel, and nothing secret was installed.

HOW, WITHOUT A KEY. `openai-compatible` is the one provider whose factory
requires a baseUrl and no apiKey. And this repository already proves worker
behaviour by replaying a recorded stream instead of calling a model - that is
what TRIOS_REPLAY_CASSETTE does on the Mac. So a route inside the container
speaks OpenAI chat-completions and streams a scripted reply, and the worker
provider points at it over loopback.

Guarded like everything else: the in-container client presents THIS SERVER'S OWN
token as its apiKey, which openai-compatible puts in the Authorization header.
The same door, used by the process itself, not a new one.

Off unless TRIOS_QUEEN_REHEARSAL is set, and never the silent fallback where a
real key exists. A hive that quietly rehearses instead of working is worse than
one that stops, because it reports success.

WHAT RAN:

  16:13:06  Queen dispatch  issue=1244 branch="queen-1244" started=true
                            cut from feat/queen-supervisor
  16:16:32  Queen rehearsal turn  model="rehearsal"
  16:16:32  Queen worker turn finished  conversationId=f20e33b7-...  issue=1244

Lease taken, stalled dispatches reaped, board read from Postgres, 40 candidates
fetched, queend chose browseros-ai#1244 by its own declared boundary, worktree cut on its
own branch, turn opened, stream consumed, dispatch recorded finished.

SAY PLAINLY WHAT IT IS NOT: no bee thought. The reply is scripted; nothing read
the issue or wrote code. What is proven is the plumbing - which is the part that
had no evidence.

TWO DEFECTS FOUND BY MAKING IT RUN.

`userWorkingDir`, not `workingDirectory`. The schema names it the first way and
ignores unknown keys, so the wrong name was accepted in silence and every bee
would have worked in the shared checkout while its branch lived in a worktree -
edits and branch in different trees, the exact failure a worktree prevents.

git ran as root. The image splits uids and the entrypoint says "git runs as bee;
root does not enter the checkout". Mine did:

  fatal: detected dubious ownership in repository at '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/workspace/BrowserOS'

The tempting fix is safe.directory and it is the wrong one: it tells git to stop
minding precisely what the uid split enforces. Dropping to bee through the same
helper every agent shell command uses keeps the split and satisfies git for the
real reason.

PARALLEL BEES UNDER THE QUEEN:

  browseros-ai#1240 -> started   rings/SR-02/ChatViewModel.swift, free
  browseros-ai#1216 -> started   docs/queen-choice.md, free
  browseros-ai#1176 -> REFUSED   QueenLocalisation.swift held by trios#1174
  browseros-ai#1286 -> REFUSED   a task already exists for it (cancelled)

Two started, two refused for different, issue-specific, correct reasons. Three
worktrees on three branches, all owned by bee.

HOW MANY IN PARALLEL - three numbers, and only the smallest matters:

  container capacity   dozens (96 concurrent, no degradation; 200/200 in 3.1s)
  the Queen's policy   4  (QueenDelegationPolicy.maximumConcurrentWorkers)
  what actually binds  boundaries - browseros-ai#1174's parked review holds
                       QueenLocalisation.swift against four issues at once

The hardware was never the limit. It is the boundary ledger and a review with no
verdict.

Still absent, and still not mine to install: a real provider key. The rehearsal
proves the plumbing, not the thinking.
Vasilev Dmitrii (gHashTag) added a commit to gHashTag/BrowserOS that referenced this pull request Aug 31, 2026
…s saying it is absent

The operator set ZAI_API_KEY. The first round after that failed:

  chat answered 500: {"message":"z.ai provider requires apiKey"}

The key was in the deployment. /chat resolves a provider from what the CALLER
supplies - its usual caller is an app on a laptop holding its own credentials -
and the server does not read its own environment for one. resolveWorkerProvider
found the key and then did not hand it over, so the round cleared the credential
precheck and died at the chat route with the exact sentence that precheck exists
to prevent. A key never passed reads identically to a key that is not there.

With the key travelling alongside the choice, a real bee ran:

  выбрана : 1244 ['BR-OUTPUT/QueenTabView.swift']
  started : True   reused an existing worktree; zai/glm-4.6

  e52f41ad Trinity Bee <bee@trinity.local>
  feat: seal QueenTabView.swift against embedded-trinity-queen-ui spec
   trios/.trinity/seals/QueenTabView.json | 16 ++++++++++++
   trios/BR-OUTPUT/QueenTabView.swift     | 30 +++++++++++++---

Chosen by the Queen, briefed by the Queen, real model, container, no laptop.

AND IT WROTE OUTSIDE ITS BOUNDARY. browseros-ai#1244 declares one path,
BR-OUTPUT/QueenTabView.swift; the commit carries two. The seal JSON is
defensible work and outside the boundary all the same.

Nothing noticed - no queen.observer.outOfBounds in the log. That observer's
cassette is one of the four that has never run since wave 069, so the first real
cloud bee did precisely the thing those cassettes exist to catch, and the
catching is the part still broken. The loop runs; its supervision has a measured
hole.
Vasilev Dmitrii (gHashTag) added a commit to gHashTag/BrowserOS that referenced this pull request Aug 31, 2026
…ased the issue

"Почему running 0" has an innocent answer and the digging found a real one.

RUNNING 0 IS NOT A FAULT. The tick fires every 30 minutes, a turn takes about
ten, so most of any half hour has nothing running. The last round was 16 minutes
before the question.

THE REAL FAULT was on the bee's own branch:

  a994a8b  12 minutes ago  sixth verification record for browseros-ai#1244 - all checks hold
  137dad0  38 minutes ago  fifth verification record for browseros-ai#1244 - all checks hold
  5b7f473  65 minutes ago  fourth verification record for browseros-ai#1244 - all checks hold
  948449b   2 hours ago    third verification record for browseros-ai#1244 - all checks hold

The work was finished hours ago. Every thirty minutes the Queen chose browseros-ai#1244
again, a bee arrived, found the job already done, verified it, and committed a
record saying so. Six times, six model turns, on one issue.

Because finishing RELEASED it. The in-flight query was `finished_at IS NULL`, so
the moment a turn ended the issue was choosable again - and it was still open on
GitHub, because nothing lands a bee's branch. A reaped dispatch must release its
issue (its container died, nothing was finished). A finished one must NOT: it
holds until somebody judges it, exactly as awaitingReview does on the Mac. The
board shows it in review now rather than letting it vanish and reappear.

Verified after the change: browseros-ai#1244 moved out of backlog into review. Round seven
will not take it.

THE BLOCKERS, measured rather than guessed:

  24 of 27 backlog issues declare no boundary, so the Queen refuses them - she
  cannot reserve files for a task that has not said which files it touches. Her
  real choosable pool is THREE, not 27.

  browseros-ai#1174 sits in review holding rings/SR-00/QueenLocalisation.swift against browseros-ai#1176
  and browseros-ai#1175. awaitingReview is not terminal. Only a verdict frees it.

MAGAZINE HEADLINES, as asked: a display scale (clamp 2.6-5.2rem, -0.045em
tracking, 0.95 leading), a letterspaced kicker above and a hairline rule under,
on all four pages.

AND THE FORM WOULD NOT HIDE. `el.hidden` was true and the box stayed on screen:
the hidden attribute is a UA rule and loses to any author display rule, so
.auth{display:flex} beat it. Measured - hidden=true, computed display=flex, 57
cards behind it. Every page now carries [hidden]{display:none !important}.

FOURTH BACKTICK OF THE DAY, in the CSS comment explaining that fix, inside a
template literal - 16 typecheck errors across four files. The SQL-only gate let
it through, so it now covers page shells too. Its first version then reported
three closing backticks as offences, because a shell closes at the end of the
last markup line rather than on a line of its own; that shape is pinned by a
test so the noise cannot come back.

typecheck 0, 280 tests across 27 files.
Vasilev Dmitrii (gHashTag) added a commit to gHashTag/BrowserOS that referenced this pull request Sep 1, 2026
Three bees finished real work on 2026-08-31 - 38 989, 56 258 and 12 989
characters of transcript across browseros-ai#1244, browseros-ai#1240 and browseros-ai#1216 - and all three were
written off with the same sentence:

  the task has no acceptance criteria, so there is nothing to judge it
  against - it can only be abandoned or accepted on faith

The policy is right and is unchanged. The criteria were in the issues. Two
things stopped them arriving, and both were mine.

`criteriaFromIssue` knew four headings: "Готово, когда", "acceptance criteria",
"done when", "готово когда". `## Success Criteria` - the heading the spec rule
I shipped a day earlier REQUIRES - was not among them. So an issue written
exactly to the rule yielded zero criteria. A rule and its only reader,
introduced a day apart and never introduced to each other.

And the cloud tick never asked. It sent `totalCriteria: verdicts.length`, which
is circular: a bee that reported nothing was judged against nothing, so the
policy saw zero criteria and escalated - to the operator, who had just said in
plain terms that the Queen must not wait on their review. The loop could start
work and structurally could not finish any.

Now: one parser in `QueenSpecQuality` (`QueenTaskSpec` delegates to it rather
than keeping a second list), carrying `## Success Criteria`, the older Russian
headings, and an `FR-001`-style fallback - an obligation the author wrote IS a
criterion. `queend` returns the criteria with each spec verdict, so TypeScript
never reimplements the rule. They are put in the brief, numbered, under "What
you will be judged by"; stored on the dispatch row, so an issue edited
mid-flight cannot change the contract a running bee was given; and read back at
review time.

When an issue names none, the bee states the criteria it will be judged by
before working, and `criteria_source` records that they are the bee's and not
the author's - so the board can say so rather than passing them off as a
contract.

Also here: `ensureQueenColumns`, because the queen tables were created by hand
against the live database and any other environment would have the code without
the columns. Columns only - a missing TABLE stays loud.

The file header claimed "nothing asks it a review question". That was the
second stale not-does claim on this file, so it is the file's pattern rather
than an accident, and it is noted as such.

Still not done, and stated rather than implied: `sendBack` is recorded and
nothing reopens the worker on the unmet criteria.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Vasilev Dmitrii (gHashTag) added a commit to gHashTag/BrowserOS that referenced this pull request Sep 1, 2026
…them busy

The operator asked why only one bee works. The refusal said

  4 workers already running (limit 4)

and exactly one bee existed. The other three had finished, been judged, and
were waiting on a person - and the board called all four `running`.

Two lines did it. The in-flight query said

  AND (finished_at IS NULL OR outcome NOT LIKE 'reaped%')

whose comment above promises "still running, OR finished with work that nobody
has judged". The condition never mentions a judgement: any finished, unreaped
dispatch matches forever. And every row it returned was put on the board with
`state: 'running'` regardless. So a task that was done in the morning was still
occupying a worker slot at midnight, and would next week.

A finished dispatch is now `awaitingReview`, which the policy already knows how
to handle: `canStartAnother` does not count it, it still blocks its own issue
from being handed out twice - that guard is why browseros-ai#1244 collected six duplicate
"verification record" commits in one afternoon - and `stillHoldsBoundary`
expires its file claim after 48 hours instead of never. Its clock starts at
finished_at, not at dispatch, or a long task would expire the moment it ended.
Its provider key is released too: a bee that has stopped is not spending it.

The `coalesce` matters: outcome is NULL while a bee runs, and
NULL NOT LIKE 'reaped%' is NULL, which drops the row. The old clause escaped
that only by ORing on finished_at.

AND THE GATE FOR THIS WAS GREEN. `sql-template-literals.test.ts` exists because
a backtick inside a SQL template literal has taken this server down twice. I
wrote a sixth one into the comment above - inside the very clause being fixed -
and the gate passed eight checks over it. It entered a block only when the
opening backtick was the LAST character on its line, and believed a literal was
SQL only on CREATE/ALTER/INSERT/UPDATE/DELETE. Every query in the supervisor is
`pool.query(` newline backtick-SELECT: opened inline, and a read. The gate had
never read one of them.

Both shapes are covered now, and the widening was proven by reintroducing the
real backtick and watching it name the line. Keeping both scanners is
deliberate: the rewrite that first caught the inline form had quietly dropped
the assignment form, which is the shape of the original outage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Vasilev Dmitrii (gHashTag) added a commit to gHashTag/BrowserOS that referenced this pull request Sep 1, 2026
The chooser skipped any issue with a task recorded against it in ANY state -
"a task already exists for it". Its own doc comment listed the states it meant:
queued, running, awaitingReview, rejected, accepted, merged. `cancelled` and
`failed` were never in that list and were excluded anyway.

So an abandoned attempt, or one that failed and therefore must be retried,
removed its issue from the board permanently. Measured on the live board this
afternoon: of 40 candidates, SIX were unreachable for this reason - browseros-ai#1127,
browseros-ai#1147, browseros-ai#1173 and browseros-ai#1286 cancelled, browseros-ai#1111 and browseros-ai#1133 failed - every one still open
on GitHub, none of them ever choosable again. A failure is the state that most
plainly means "do this again", and it was being read as "never do this".

`QueenDelegationPolicy.claimOnIssue` now answers in three ways instead of one:
live (queued, running, awaitingReview, rejected - someone has it or is expected
back), done (accepted, merged - choosing it again would redo landed work, which
is how browseros-ai#1244 collected six duplicate verification commits in an afternoon), or
free. Only free is choosable, and cancelled and failed are free.

It reads EVERY task for the issue rather than the first one the registry
happens to list, so a live retry over an old failure is claimed by the retry.

The refusals say which of the three now, rather than one sentence for all six
states. On the live board that turned "a task already exists for it" into "the
work already landed (accepted) - the issue is open and nobody closed it", which
is a different problem and now says so.

Gated by `queend-choose.test.ts`, which drives the real binary because
queen-core has no test target and XCTest does not link under the
CommandLineTools toolchain. Proven both ways: 6 of its 7 checks fail against
the old chooser, all 7 pass against this one. Its skip when the binary is
absent is stated in the file rather than left to be discovered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Vasilev Dmitrii (gHashTag) added a commit to gHashTag/BrowserOS that referenced this pull request Sep 1, 2026
A skeptic broke my own last commit on a live harness, and it was right.

The history archive added to `recordDispatch` put one more statement in front
of the upsert. `startTurn` fired the stream reader with `void drain(...)` and
returned; the caller wrote the row afterwards. So the two raced - and with the
archive in front, the reader could win. Measured on a real stream with a 4 ms
pool, a turn whose frames are all NOISE does no database work of its own and
arrived in this order:

  INSERT INTO queen_dispatch_history
  UPDATE queen_dispatch      <- the ending
  INSERT INTO queen_dispatch <- the row it belongs to

The ending matched zero rows. An UPDATE that changes nothing does not throw, so
the `try` around it saw success, and the upsert then wrote started=true with
finished_at NULL. A bee that had already stopped looked like it was running and
held its files until the 120-minute reaper - which is the exact phantom the
previous commit added logging to catch, arriving by a path with no database
failure in it at all.

Three changes, and the first is the real one:

`startTurn` hands the reader back instead of firing it. `dispatchBee` calls
`turn.beginDrain()` AFTER the row is written. Everything that reads the bee's
output eventually writes to that row, and a writer that can outrun the row's
creation is a writer that silently updates nothing.

`closeDispatch` checks rowCount. Zero rows is now an error line naming the
issue and the conversation, because "I wrote the ending" and "I matched no row"
were the same answer.

`finishDispatch` takes the conversation and puts it in the WHERE. Keyed by
issue alone, a stream from a previous attempt that finishes late closes the
CURRENT attempt's row and drives its token counts straight through the COALESCE
that exists to protect a price. Reaping releases an issue for retry while the
old container's stream may still be alive, so two turns overlapping on one
issue is routine, not exotic - it is the six-turns-on-browseros-ai#1244 shape seen from the
other side. The parameter is optional so the reaper, which legitimately closes
a row whose conversation is gone, still can.

Proven by deleting the rowCount branch and watching 'says so when the ending
matched no row at all' go red, then restoring it. 402 tests green, typecheck 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant