Skip to content

fix(server): report delegated tasks stopped by a server restart - #13345

Open
saphid wants to merge 2 commits into
pingdotgg:t3code/codex-turn-mappingfrom
saphid:fix/v2-subagent-recovery-reports
Open

saphid wants to merge 2 commits into
pingdotgg:t3code/codex-turn-mappingfrom
saphid:fix/v2-subagent-recovery-reports

Conversation

@saphid

@saphid saphid commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

What Changed

When the server restarts, a parent now hears about delegated tasks the restart stopped as soon as the server is back up, not at the next restart.

On startup, runtime recovery cancels the runs of delegated children that were still working, and any parent completion wake that was running. Those cancellations use runtime-reconcile commands, which the live terminal-run reactor deliberately ignores. The orchestrator did recover unreported child results and interrupted completion wakes, but it did so while being constructed, before runtime recovery ran. Anything recovery cancelled was left until the following server start.

This change moves both recovery passes, child results and completion wakes, into one recoverDelegatedTaskReports step. It runs once, in the startup recover phase, right after runtime recovery and before the effect worker starts. Results go through the existing completion delivery.

Restart continuation ("continue threads after update") is respected:

  • A child whose restart continuation is still pending in the outbox is not reported. If the continuation starts a run, the continued run reports as usual.
  • If the continuation is skipped (for example, the setting is off or newer work exists), or it fails for good, the executor publishes the cancelled result then. continueRestartedRun now returns whether it dispatched a run.
  • Because nothing runs during construction any more, a child with a pending continuation is never reported cancelled first, including after a crash between recovery and the continuation.
  • Cancelled completion wakes are finalized without deferral. Their claimed tasks go back to pending and get a follow-up delivery. A restart continuation of a wake does not carry its delivery ownership, so deferring would strand those tasks.

Why

In our database, 11 delegated children were cancelled across two parents at 23:42 UTC. Their parents only received the results at 03:47, when the server next started. That's four hours in which each parent assumed its children were still working.

This is one of a few focused fixes to make sure a parent learns when a child stops. #13938 fixed results that were never delivered, and #13343 covers children blocked on a question or approval.

Tests

  • DelegatedCompletionDelivery.test.ts, "recovers a child that runtime recovery cancelled, deferring pending continuations": nothing is reported before the pass runs; a child with a pending continuation is skipped; otherwise its result is published once, and a second pass publishes nothing. If the deferral is removed, the test fails.
  • DelegatedCompletionDelivery.test.ts, "re-offers a parent wake that startup recovery cancelled with a sibling": a running wake and a running sibling are both cancelled by recovery. The pass publishes the sibling's result and reserves one follow-up delivery carrying both tasks. If the wake pass is dropped, the test fails.
  • RestartContinuation.test.ts, "reconciles a delegated child when its restart continuation will not run": nothing happens when the continuation starts or will retry; the result is published when it is skipped or fails for good.
  • The earlier head passed 138 tests in serverRuntimeStartup, ProviderRuntimeRecoveryService (and its regression file), RestartContinuation, EffectWorker, DelegatedCompletionDelivery, AgentAwarenessRelay, ThreadManagementService, OrchestratorMcpToolkit.integration, SteeringCompletion.integration, ThreadDeletion, and runtimeLayer. The latest recovery-error fix passes vp test run apps/server/src/orchestration-v2/ProviderRuntimeRecoveryService.test.ts apps/server/src/orchestration-v2/DelegatedCompletionDelivery.test.ts (19 tests) and pnpm run typecheck in apps/server.

Limits:

  • No test rebuilds the orchestrator over saved state. That construction no longer recovers anything is a structural change, verified by reading the code.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (no UI change)
  • I included a video for animation/interaction changes (no UI change)

This change was made by Claude Opus 5.5 with Claude Code and reviewed by GPT-6 Astra with Codex.

🤖 Generated with Claude Code

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Sep 24, 2026
Comment thread apps/server/src/serverRuntimeStartup.ts
@macroscopeapp

macroscopeapp Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production fix changes startup recovery ordering and delegated-task orchestration across multiple shared runtime components. It also gates child-result publication and parent wake processing on restart-continuation state, creating a broader runtime impact than a small self-contained bug fix.

You can add or adjust custom eligibility rules. Learn more.

github-actions Bot and others added 2 commits September 28, 2026 09:23
When the server restarts, startup recovery cancels the runs of delegated
children that were still working, and any parent completion wake that
was running. Those cancellations use runtime-reconcile commands, which
the live terminal-run reactor ignores, and the orchestrator recovered
child results and completion wakes while it was constructed, before
startup recovery ran. So a parent only learned that its child had
stopped at the next server start, which could be hours later.

Move both recovery passes to run once startup recovery has finished,
and publish results through the existing completion delivery. A child
with a restart continuation still pending is left to that continuation:
if the continuation starts a run, the continued run reports as usual;
if it is skipped or fails for good, the cancelled result is published
then. Recovering after startup recovery also means a child with a
pending continuation is never reported cancelled first.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L 100-499 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant