Skip to content

fix(server): startup recovery settles native subagent threads - #13619

Merged
juliusmarminge merged 2 commits into
t3code/codex-turn-mappingfrom
v2/subagent-recovery
Sep 25, 2026
Merged

juliusmarminge merged 2 commits into
t3code/codex-turn-mappingfrom
v2/subagent-recovery

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

If the server died while a provider-native subagent (for example a Claude Agent-tool subagent) was still working, its child thread stayed running forever after restart. Since #13614, clients show a runless running root turn as working, so the child would read "Working for 3h".

Why. A native subagent's child thread has no runs. Its work is a single runless root_turn whose status follows the subagent, and only the provider process can move it. Startup recovery already cancels the parent's stale subagent turn item, entity, and node, but it never visited the child thread: none of the getRecoveryThreadIds("runtime") candidates match a thread whose only live state is a runless root turn. The recovery projection's node query also only loaded nodes owned by live runs or background items.

What changed

  • ProjectionStore.getRecoveryThreadIds("runtime"): adds child threads of subagents that still have an active runless root turn. The lookup goes through the subagents.child_thread_id index plus the nodes(thread_id, run_id) index, so settled histories are still never scanned.
  • getRuntimeRecoveryProjection: loads active runless root turns.
  • ProviderRuntimeRecoveryService: cancels them (status cancelled, completedAt = recovery time), the same way it settles other process-bound nodes. Shutdown reconciliation takes the same path.

Independent of #13614; either can land first.

Verification

  • New FoundationPersistence.test.ts case against the real SQLite projection store and event sink. The parent has a settled run with a running native subagent (item, entity, node), and the child has a runless running root turn. After recover, the parent subagent is cancelled, the child root turn is cancelled with a completedAt, and the child is no longer a recovery candidate. Without the source change it fails: the child is never selected for recovery.
  • cd apps/server && vp test run src/orchestration-v2/FoundationPersistence.test.ts src/orchestration-v2/ProjectionRecovery.test.ts src/orchestration-v2/ProviderRuntimeRecoveryService.test.ts src/orchestration-v2/ProviderRuntimeRecoveryService.regression.test.ts src/orchestration-v2/RestartContinuation.test.ts: 63 passed.
  • vp exec tsc --noEmit -p . (server): no errors. vp lint on touched files: clean. vp run knip:check: clean.
  • EXPLAIN QUERY PLAN on the new candidate branch: SEARCH subagent USING COVERING INDEX ..._subagents_child_thread_idx then SEARCH node ... USING INDEX ..._nodes_thread_run_idx (thread_id=? AND run_id=?).
  • Not run: a real crash mid-subagent against a live provider, or repo-wide checks.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

A provider-native subagent's child thread has no runs; its work is a
runless root turn that only the provider process can settle. When the
server died mid-subagent, startup recovery cancelled the parent's
subagent item, entity, and node, but never looked at the child thread:
none of the runtime recovery candidates matched a thread whose only live
state is a runless root turn. The child stayed "running" forever, which
clients now show as working.

Select child threads with an active runless root turn for runtime
recovery, load those root turns into the recovery projection, and cancel
them alongside the other process-bound work.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 25, 2026
Comment thread apps/server/src/orchestration-v2/ProviderRuntimeRecoveryService.ts
@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.8 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: 7d481c0 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@macroscopeapp

macroscopeapp Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 7d481c0

Macroscope's review found this PR approvable — This is a focused startup/shutdown recovery fix that discovers runless native-subagent child roots and terminalizes their live items, with an integration test covering streaming reasoning cleanup. The supplied Medium concern about omitted runless reasoning items appears addressed in the final query and test; no product defaults, schema, security, or static-analysis settings are changed.

Notes:

  • This verdict was updated automatically after the outstanding correctness findings were resolved. Macroscope did not re-review the code.

You can add or adjust custom eligibility rules. Learn more.

Startup recovery cancelled a native subagent's runless root turn but left
items under it running, such as Claude's streaming "Subagent progress"
reasoning item, so the recovered child still showed live work. Load the
active runless root turn's items into the recovery projection and cancel
them with the root turn.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
});
}
}
// A provider-native subagent thread has no runs: its work is a runless

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new runless-root-turn recovery path needs a focused test. Existing recovery tests cover runless background items and run-backed nodes, but not discovery and cancellation of a runless child root_turn with a streaming reasoning item. Could you add a test using the real projection store and test layers for external services that verifies the child is selected, the node and item are cancelled, and streaming becomes false?

Posted via Macroscope — Effect Service Conventions

@juliusmarminge
juliusmarminge merged commit 6c39f90 into t3code/codex-turn-mapping Sep 25, 2026
24 checks passed
@juliusmarminge
juliusmarminge deleted the v2/subagent-recovery branch September 25, 2026 19:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M 30-99 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant