Skip to content

fix(server): parents keep getting woken for every delegated task that finishes - #13938

Merged
juliusmarminge merged 1 commit into
t3code/codex-turn-mappingfrom
v2/delegated-completion-no-cap
Sep 27, 2026
Merged

juliusmarminge merged 1 commit into
t3code/codex-turn-mappingfrom
v2/delegated-completion-no-cap

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Closes #12285

A parent that fans out delegated tasks with delegate_task stopped being woken after two completion deliveries. The rest of its children kept finishing, but their results stayed pending with no wake, and the parent sat "waiting" until the user typed. In the reported DB, both originating runs show delegatedCompletion.settledDeliveryCount: 2 and delivery: null: two notices got through, and 10 (then 5) completions were silent.

Cause

#5311 coalesced sibling completions into one durable delivery per parent-run cohort, which was the right goal. But Orchestrator.ts also refused to reserve anything once settledDeliveryCount >= 2, in three places: planning a new delivery, reserving a successor when a delivery settles, and batching on mailbox acceptance. So the cohort was hard-capped at one delivery plus one successor.

Fix

  • Removed the cap. When a delivery settles and terminal results are still pending, reserve one successor carrying all of them.
  • Everything else from fix(orchestrator): Prevent redundant delegated completion turns #5311 stays as it was:
    • Siblings that finish before a delivery starts join it.
    • Results that land while a delivery is queued or running wait for it to settle, then go in the next batch. A cohort never has more than one outstanding delivery.
    • Acknowledgement through task_status and direct-child reads is unchanged.
    • Stop, Queue Remove, archive and delete still act as barriers, and recovery and replay are untouched.
  • Why no replacement limit: each child goes from running to pending at most once, so successors are bounded by the cohort's children. There is no recursive per-child re-arming. A loop where each delivery makes the parent spawn more children isn't possible within one cohort either: delegate_task attaches new children to the currently active run, and a delivery run is a new run with its own cohort. Any bound would reintroduce silently dropped wakes, so I didn't add one.
  • settledDeliveryCount had no other reader, so I dropped it from OrchestrationV2DelegatedCompletionCohort. Persisted rows that still carry it decode fine, because the schema ignores excess properties by default.

Verification

  • Late-child test in OrchestratorMcpToolkit.integration.test.ts: third and fourth late children now finish while the successor runs. They stay pending behind the one outstanding reservation, then go out together in a third delivery. All four children end delivered, with one continuation offer per delivery (3).
  • New 12-child scenario, same real orchestrator harness:
    • 2 children finish, the first delivery starts, 5 finish during it, 3 join the queued successor, and 2 finish during the successor.
    • Result: exactly 3 batched deliveries covering all 12 children. Every child ends delivered, with no pending results left, and at most one delivery outstanding at each checkpoint.
  • Negative check: with the base Orchestrator.ts and contract restored, the updated test fails, timing out waiting for the batched third delivery.
  • Tests run:
    • vp test run src/mcp/OrchestratorMcpToolkit.integration.test.ts: 2/2 pass.
    • DelegatedCompletionDelivery, ProviderContinuationService, ProjectionStore, QueuedRunOrder, ThreadDeletion, CheckpointCaptureService, SteeringCompletion.integration, SubagentProjection, ProjectionRecovery, RunCompletionReads, TurnStartReads and OrchestratorMcpService tests: 79/79 pass.
    • OrchestratorReplayFixtures.integration.test.ts: 96/96 pass.
  • Static checks:
    • vp exec tsc --noEmit -p . in apps/server and packages/contracts: no errors.
    • knip (--exports, server and contracts workspaces): clean.
    • vp lint on the touched files: only the layerUnavailable unused-variable warning, which is already on the base branch.
  • Not run: repo-wide checks, a live provider run, and web or mobile. Clients don't read settledDeliveryCount.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

… finishes

A delegated-completion cohort allowed one delivery plus one successor, then
left every later child result pending with no wake. A parent that fanned out
twelve children was woken twice and then waited silently until the user typed.

When a delivery settles with results still pending, reserve one successor that
carries all of them. Siblings that finish before a delivery starts still join
it, and results that land while it is queued or running wait for the next one,
so a cohort never has more than one delivery outstanding. Each child becomes
pending at most once, so successors are bounded by the cohort's children.

The cap was the only reader of settledDeliveryCount, so the field is dropped
from the contract; older rows that carry it decode unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 27, 2026
@macroscopeapp

macroscopeapp Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 78d50b9

Macroscope's review found this PR approvable — This is a small, isolated orchestration bug fix that restores delivery of late delegated-task results without changing unrelated execution paths. Extensive integration coverage verifies batching, one-outstanding-delivery behavior, cleanup barriers, and complete delivery across a 12-child fan-out.

You can add or adjust custom eligibility rules. Learn more.

@github-actions

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.1 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 1 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.8 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: 78d50b9 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@juliusmarminge
juliusmarminge merged commit b39e62d into t3code/codex-turn-mapping Sep 27, 2026
24 of 25 checks passed
@juliusmarminge
juliusmarminge deleted the v2/delegated-completion-no-cap branch September 27, 2026 21:43
SoulSniper-V2 pushed a commit to SoulSniper-V2/t3code that referenced this pull request Sep 27, 2026
… finishes (pingdotgg#13938)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M 30-99 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant