fix(nightly-v2): keep startup recovery running past undecodable thread projections - #89
Merged
Merged
Conversation
…ections Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
saphid
marked this pull request as ready for review
September 23, 2026 04:08
2 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What Changed
Pin the V2 startup recovery fix into Fork Nightly. When one thread's persisted projection cannot be decoded, runtime recovery now logs a warning and skips that thread. The server still starts. Storage failures still stop startup, as before.
Fork
maincarries release configuration, not V2 source. This PR adds one source commit after the existing 14 patches.Why
On 2026-09-23 the standalone V2 service (
0.0.42-nightly-v2.20260918) crash-looped every 5 to 10 seconds, and the desktop showed "AUS-M5P-AS is reconnecting" until the database was repaired by hand. The boot log showed the same failure on every start:Ten delegated-task turn items in
statev2.sqlitehadfailure.class: "usage_limit"(Claude 429s). A newer build wrote that value; the running 0918 build's schema does not accept it.reconcile("startup")read every recovery candidate withgetRuntimeRecoveryProjectionand turned the first decode failure into a fatal startup error. The service launcher restarted the server into the same failure every time.Current Fork Nightly already accepts
usage_limit, so that particular value no longer breaks it. The same failure returns whenever one build reads a value from a newer build that shares its state directory, or finds any other undecodable row. That can happen with the standalone service and a desktop-bundled server running different versions, or after a downgrade. One unreadable thread should cost that thread's recovery, not the whole server. The store already tracks unreadable threads throughgetUnreadableThreadIds. Recovery now treats them the same way.The predicate walks the
ProjectionStoreReadErrorcause chain and matches only aSchema.SchemaErrorat its root. Undecodable child rows and undecodable thread rows are both skipped. SQL and other storage errors are not matched and still fail startup.prepareForShutdownuses the same skip, so shutdown cannot fail on the same row either.A skipped thread's nonterminal runs, runtime requests and sessions are left as they were and are not reconciled until its row is repaired or a build that understands the value reads it. Durable effects are still reconciled by
outbox.reconcileAfterProcessLoss, which only touches the outbox table. Restart continuation needs a cancelled source run, so a skipped running source is not replayed.Verification
ProjectionRecovery.test.ts, in-memory SQLite, realProjectionStore) writes a run whose status only a newer build would produce. On base593a99f928it fails with the production chain:ProviderRuntimeRecoveryErrorread-projections, caused bySchemaError: Expected "preparing" | "queued" | .... With the fix, recovery completes and reconciles both healthy threads, including one ordered after the skipped thread. Replacing thecontinuewithbreakmakes the test fail.ProjectionStoreReadErrorcaused by a storage error (SQLITE_IOERR) still fails recovery withread-projectionsand the thread id.vp test run apps/server/src/orchestration-v2/onea259e33fe: 90 files passed, 3 skipped; 1321 tests passed, 6 skipped. Head843307eb20changes only the new test; the three recovery test files rerun on it: 19 passed.vp exec tsc --noEmit -p apps/server/tsconfig.json: passed. Focusedvp lintandvp fmt --check: passed.t3code/codex-turn-mappingat060756de5adcda82fedace42d8c47b449bc1097bapplied cleanly. The resulting recovery sources match the tested tree byte for byte; the only other differences are 4 generatedpackage.jsonversions.node --test .github/scripts/downstream-nightly.test.mjs: 22 passed. The V2 manifest parses with 15 patches.ea259e33fethrough Codex CLI: approve, no blocking findings. It confirmed the predicate matches how the store wraps errors (one wrapper for thread rows, two for child rows, with JSON parsing also surfacing asSchemaError), that SQL errors stay fatal, and that no other startup path aborts on the same row. Its only finding (P3) was that the test could not tellcontinuefrombreak. Head843307eb20adds the later healthy thread to fix that. A delta review of843307eb20also approved, with no findings. It confirmed that recovery orders candidates byupdated_at, thread_id, so the new thread really does come after the skipped one.usage_limittoprovider_error, originals saved). The regression test reproduces the same error chain instead.ci.ymlrun since 2026-09-23 02:44Z, includingmain. PR fix(nightly-v2): unblock releases and defer optional features #86 merged with the same checks cancelled. The server tests above were run locally instead.Checklist
Implementation: Claude Opus 5.5 in Claude Code. Independent review: GPT-6 Astra in Codex CLI, high effort.
🤖 Generated with Claude Code