fix(server): startup recovery settles native subagent threads - #13619
Conversation
A provider-native subagent's child thread has no runs; its work is a runless root turn that only the provider process can settle. When the server died mid-subagent, startup recovery cancelled the parent's subagent item, entity, and node, but never looked at the child thread: none of the runtime recovery candidates matched a thread whose only live state is a runless root turn. The child stayed "running" forever, which clients now show as working. Select child threads with an active runless root turn for runtime recovery, load those root turns into the recovery projection, and cancel them alongside the other process-bound work. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
ApprovabilityVerdict: Approved at Macroscope's review found this PR approvable — This is a focused startup/shutdown recovery fix that discovers runless native-subagent child roots and terminalizes their live items, with an integration test covering streaming reasoning cleanup. The supplied Medium concern about omitted runless reasoning items appears addressed in the final query and test; no product defaults, schema, security, or static-analysis settings are changed. Notes:
You can add or adjust custom eligibility rules. Learn more. |
Startup recovery cancelled a native subagent's runless root turn but left items under it running, such as Claude's streaming "Subagent progress" reasoning item, so the recovered child still showed live work. Load the active runless root turn's items into the recovery projection and cancel them with the root turn. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| }); | ||
| } | ||
| } | ||
| // A provider-native subagent thread has no runs: its work is a runless |
There was a problem hiding this comment.
The new runless-root-turn recovery path needs a focused test. Existing recovery tests cover runless background items and run-backed nodes, but not discovery and cancellation of a runless child root_turn with a streaming reasoning item. Could you add a test using the real projection store and test layers for external services that verifies the child is selected, the node and item are cancelled, and streaming becomes false?
Posted via Macroscope — Effect Service Conventions
If the server died while a provider-native subagent (for example a Claude Agent-tool subagent) was still working, its child thread stayed
runningforever after restart. Since #13614, clients show a runless running root turn as working, so the child would read "Working for 3h".Why. A native subagent's child thread has no runs. Its work is a single runless
root_turnwhose status follows the subagent, and only the provider process can move it. Startup recovery already cancels the parent's stale subagent turn item, entity, and node, but it never visited the child thread: none of thegetRecoveryThreadIds("runtime")candidates match a thread whose only live state is a runless root turn. The recovery projection's node query also only loaded nodes owned by live runs or background items.What changed
ProjectionStore.getRecoveryThreadIds("runtime"): adds child threads of subagents that still have an active runless root turn. The lookup goes through thesubagents.child_thread_idindex plus thenodes(thread_id, run_id)index, so settled histories are still never scanned.getRuntimeRecoveryProjection: loads active runless root turns.ProviderRuntimeRecoveryService: cancels them (statuscancelled,completedAt= recovery time), the same way it settles other process-bound nodes. Shutdown reconciliation takes the same path.Independent of #13614; either can land first.
Verification
FoundationPersistence.test.tscase against the real SQLite projection store and event sink. The parent has a settled run with a running native subagent (item, entity, node), and the child has a runless running root turn. Afterrecover, the parent subagent is cancelled, the child root turn iscancelledwith acompletedAt, and the child is no longer a recovery candidate. Without the source change it fails: the child is never selected for recovery.cd apps/server && vp test run src/orchestration-v2/FoundationPersistence.test.ts src/orchestration-v2/ProjectionRecovery.test.ts src/orchestration-v2/ProviderRuntimeRecoveryService.test.ts src/orchestration-v2/ProviderRuntimeRecoveryService.regression.test.ts src/orchestration-v2/RestartContinuation.test.ts: 63 passed.vp exec tsc --noEmit -p .(server): no errors.vp linton touched files: clean.vp run knip:check: clean.EXPLAIN QUERY PLANon the new candidate branch:SEARCH subagent USING COVERING INDEX ..._subagents_child_thread_idxthenSEARCH node ... USING INDEX ..._nodes_thread_run_idx (thread_id=? AND run_id=?).Model: Claude Opus 5.5 (Claude Code)
🤖 Generated with Claude Code