Skip to content

fix(nightly-v2): keep startup recovery running past undecodable thread projections - #89

Merged
saphid merged 1 commit into
mainfrom
fix/ov2-pin-recovery-undecodable-thread
Sep 23, 2026
Merged

saphid merged 1 commit into
mainfrom
fix/ov2-pin-recovery-undecodable-thread

Conversation

@saphid

@saphid saphid commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

What Changed

Pin the V2 startup recovery fix into Fork Nightly. When one thread's persisted projection cannot be decoded, runtime recovery now logs a warning and skips that thread. The server still starts. Storage failures still stop startup, as before.

Fork main carries release configuration, not V2 source. This PR adds one source commit after the existing 14 patches.

Why

On 2026-09-23 the standalone V2 service (0.0.42-nightly-v2.20260918) crash-looped every 5 to 10 seconds, and the desktop showed "AUS-M5P-AS is reconnecting" until the database was repaired by hand. The boot log showed the same failure on every start:

ProviderRuntimeRecoveryError: Provider runtime recovery failed during read-projections.
  ProjectionStoreReadError: Failed to read orchestration projection for thread thread:delegated-task:...foreman-worker-wscloud-C05-20260923T0200Z.
    SchemaError: Expected "provider_error" | "transport_error" | "permission_error" | "validation_error" | "unknown"
      at ["failure"]["class"]

Ten delegated-task turn items in statev2.sqlite had failure.class: "usage_limit" (Claude 429s). A newer build wrote that value; the running 0918 build's schema does not accept it. reconcile("startup") read every recovery candidate with getRuntimeRecoveryProjection and turned the first decode failure into a fatal startup error. The service launcher restarted the server into the same failure every time.

Current Fork Nightly already accepts usage_limit, so that particular value no longer breaks it. The same failure returns whenever one build reads a value from a newer build that shares its state directory, or finds any other undecodable row. That can happen with the standalone service and a desktop-bundled server running different versions, or after a downgrade. One unreadable thread should cost that thread's recovery, not the whole server. The store already tracks unreadable threads through getUnreadableThreadIds. Recovery now treats them the same way.

The predicate walks the ProjectionStoreReadError cause chain and matches only a Schema.SchemaError at its root. Undecodable child rows and undecodable thread rows are both skipped. SQL and other storage errors are not matched and still fail startup. prepareForShutdown uses the same skip, so shutdown cannot fail on the same row either.

A skipped thread's nonterminal runs, runtime requests and sessions are left as they were and are not reconciled until its row is repaired or a build that understands the value reads it. Durable effects are still reconciled by outbox.reconcileAfterProcessLoss, which only touches the outbox table. Restart continuation needs a cancelled source run, so a skipped running source is not replayed.

Verification

  • A new real-store regression test (ProjectionRecovery.test.ts, in-memory SQLite, real ProjectionStore) writes a run whose status only a newer build would produce. On base 593a99f928 it fails with the production chain: ProviderRuntimeRecoveryError read-projections, caused by SchemaError: Expected "preparing" | "queued" | .... With the fix, recovery completes and reconciles both healthy threads, including one ordered after the skipped thread. Replacing the continue with break makes the test fail.
  • A new regression test confirms a ProjectionStoreReadError caused by a storage error (SQLITE_IOERR) still fails recovery with read-projections and the thread id.
  • vp test run apps/server/src/orchestration-v2/ on ea259e33fe: 90 files passed, 3 skipped; 1321 tests passed, 6 skipped. Head 843307eb20 changes only the new test; the three recovery test files rerun on it: 19 passed.
  • Server vp exec tsc --noEmit -p apps/server/tsconfig.json: passed. Focused vp lint and vp fmt --check: passed.
  • A local replay of all 15 patches onto upstream t3code/codex-turn-mapping at 060756de5adcda82fedace42d8c47b449bc1097b applied cleanly. The resulting recovery sources match the tested tree byte for byte; the only other differences are 4 generated package.json versions. node --test .github/scripts/downstream-nightly.test.mjs: 22 passed. The V2 manifest parses with 15 patches.
  • GPT-6 Astra at high effort reviewed source ea259e33fe through Codex CLI: approve, no blocking findings. It confirmed the predicate matches how the store wraps errors (one wrapper for thread rows, two for child rows, with JSON parsing also surfacing as SchemaError), that SQL errors stay fatal, and that no other startup path aborts on the same row. Its only finding (P3) was that the test could not tell continue from break. Head 843307eb20 adds the later healthy thread to fix that. A delta review of 843307eb20 also approved, with no findings. It confirmed that recovery orders candidates by updated_at, thread_id, so the new thread really does come after the skipped one.
  • The fix has not been run against the affected production database. That database was repaired first so the live service could come back (10 rows rewritten from usage_limit to provider_error, originals saved). The regression test reproduces the same error chain instead.
  • CI: "Validate fork Nightly automation" passed on this head. Check, Test, Release Smoke and Mobile Native Static Analysis are still queued with no Blacksmith runner, like every fork ci.yml run since 2026-09-23 02:44Z, including main. PR fix(nightly-v2): unblock releases and defer optional features #86 merged with the same checks cancelled. The server tests above were run locally instead.
  • Backend startup change only; nothing in the UI changes.

Checklist

  • This PR is small and focused
  • I explained what changed and why

Implementation: Claude Opus 5.5 in Claude Code. Independent review: GPT-6 Astra in Codex CLI, high effort.

🤖 Generated with Claude Code

…ections

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XS labels Sep 23, 2026
@saphid
saphid marked this pull request as ready for review September 23, 2026 04:08
@saphid
saphid merged commit 3e55d8e into main Sep 23, 2026
13 of 17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XS vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant