Skip to content

delegate_task mode:"wait" dies at the client ~5 min ceiling and returns no taskId, causing duplicate child dispatch #11168

Description

@astarktc

Summary

delegate_task with mode: "wait" is unsurvivable for any child that runs longer than ~5 minutes when the MCP client is Node-based (Pi provider here, but this is true of any fetch/undici client). The wait is a single synchronous MCP request held open for the child's entire run; undici's default headersTimeout is 300 s, so the HTTP request dies first and the caller gets a bare fetch failed — with no taskId and no childThreadId.

The consequence is worse than a lost result: the parent agent has no handle to reconcile with, cannot call task_status, and cannot tell "the child never launched" from "the child is running fine". The only apparent recourse is re-dispatch — which spawns a second child writing into the same cwd.

Evidence (two independent occurrences, same day, same fingerprint)

parent A parent B
child dispatched 21:20:12Z 01:10:01Z
child actually ran to completed 21:55:50Z 01:15:03Z
parent's wait call failed → duplicate re-dispatch 21:25:52Z (+5m40s) 01:15:52Z (+5m51s)
duplicate child interrupted 21:27:02Z interrupted 01:18:17Z

Both retries landed at ≈5m45s after dispatch — the undici 300 s ceiling plus one agent turn. Control: a trivial mode:"wait" child (seconds long) returns a clean, complete envelope (taskId, childThreadId, status: completed, summary), so wait-mode itself is fine; only its duration tolerance is broken.

In the second case the duplicate child was smart enough to notice the first child's commit and refuse to redo the work. That was luck, not a guarantee — two concurrent writers in one working tree is the real hazard.

Where the mismatch lives

apps/server/src/mcp/OrchestratorMcpService.ts clamps the wait budget to a default of 10 minutes and a maximum of 60 minutes. Both sit well above the ~5-minute ceiling a Node fetch client can hold a response open, and nothing is emitted on the wire during the wait to keep the connection alive. The tool description also advertises waitTimedOut + "keep that taskId and read status on later task_status" — but that graceful path is only reachable when the server's own timer fires first, which by construction it cannot for the default budget.

Related: when the parent's turn was handed off/interrupted while a wait was in flight, the delegated child was interrupted too, so a child's lifetime appears coupled to the parent's in-flight tool call.

Suggested fixes (any one helps; 1+3 would close it)

  1. Emit MCP progress notifications (or SSE keep-alives) during mode: "wait". Traffic on the stream resets the client's header/body timers, which makes long waits viable and is the behavior most MCP clients expect for long-running tools.
  2. Or clamp the default/max wait below the client ceiling (e.g. default 120 s, max 240 s) and return the documented waitTimedOut: true envelope, which at least preserves the taskId.
  3. Make the taskId observable before the wait resolves — e.g. send it in an early progress notification, so a parent whose wait died can reconcile via task_status instead of re-dispatching. Documenting "on any delegate_task transport error, call t3_thread_list before retrying" would also help agents, since a server-generated clientRequestId cannot dedupe a retry (the caller never saw it).

Environment

macOS 15, Pi provider (anthropic/claude-opus-5), full-access runtime. Build is a local fork tracking the Orchestrator V2 branch (PR #2829) at 8f44bec / 0.0.40, with two unrelated local patches (Pi in the Usage dashboard; delegated-wake cap removal) — neither touches the MCP layer. The code paths cited are upstream-unmodified.

Activity

  1. juliusmarminge commented on Sep 11, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed as a real Orchestrator V2 MCP bug. The report matches the current V2 branch (t3code/codex-turn-mapping / PR #2829). This code is not on main.

    What breaks

    delegate_task with mode: "wait" creates the child, then holds the same MCP tools/call open while waitForTask polls every 50ms until the child is terminal or the server wait budget elapses:

    • default wait: 10 minutes
    • max wait: 60 minutes

    Nothing is written on the HTTP/SSE stream during that poll. A Node fetch / undici client dies first at the default 300s header/body timeout and surfaces a bare fetch failed with no tool result. taskId and childThreadId exist on the server before the wait starts, but they are only returned after the wait resolves. The caller therefore has no handle for task_status and cannot tell "never launched" from "still running". Re-dispatch is the natural next step and creates a second child in the same cwd.

    The two ~5m45s retries in the report match that ceiling plus one agent turn. Short waits still return a complete envelope, so wait-mode itself works; only long waits are unsurvivable.

    Why the documented recovery cannot fire

    The tool description and schema (from PR #7427) tell agents that timeoutMs is only the parent's wait budget, that waitTimedOut: true does not cancel the child, and that they should keep taskId and poll task_status. That envelope is only produced when the server timer wins. Against a 5-minute client ceiling and a 10-minute default, it cannot.

    clientRequestId is idempotent when the caller supplies the same value (integration tests replay delegate-claude-1 and get the same taskId). If it is omitted, the server generates a random UUID, so a retry cannot dedupe.

    t3_thread_list can list subagent children (includeSubagents defaults to true). The tool text does not tell agents to use that before retrying, and there is no task_list.

    Wait-mode also sets completionWake: "settled_only" and only upgrades to "always" on the server timeout path. A client disconnect leaves the original child running without that upgrade, so a later completion may not wake a still-active parent.

    The first child in the report ran to completed after the wait died, so the fetch timeout itself does not interrupt the child. Parent handoff/interrupt coupling looks like a separate question, not this failure mode.

    t3_thread_wait reuses the same 10/60-minute silent budget and has the same keep-alive gap.

    Related

    Suggested fix (on the V2 branch)

    1. Emit MCP progress notifications (or SSE keep-alives) during mode: "wait" so Node client timers reset.
    2. Publish taskId / childThreadId in an early progress notification so a dead wait can still be reconciled with task_status.
    3. Treat client disconnect like server timeout for the completionWake upgrade.

    Optional safety net: clamp default/max wait below typical client ceilings so the documented waitTimedOut: true envelope can actually return. Also document: on any delegate_task transport error, call t3_thread_list before re-dispatch, and pass a caller-owned clientRequestId.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 11, 2026
  3. Alb11747 commented on Oct 9, 2026

    @Alb11747

    Additional confirmed occurrence on Windows x64, T3 Code 0.0.46-nightly.20261008.2849, using the Codex provider (codex-proxy, gpt-6.1-sol). This is the same 300-second failure, with nested delegation explaining why a short triage run held the outer wait open.

    Observed timeline (2026-10-09, UTC)

    The scheduled parent called:

    {
      "target": {
        "providerInstanceId": "codex-proxy",
        "model": "gpt-6.1-sol",
        "options": { "reasoningEffort": "medium" }
      },
      "role": "review",
      "mode": "wait",
      "timeoutMs": 600000,
      "runtimeMode": "full-access",
      "clientRequestId": "fobbitmc-triage-20261009T0305",
      "task": "Read the monitor instructions and state, perform triage, and return a one-line result."
    }

    The task text above summarizes the original file-based prompt; the remaining fields are exact.

    • 03:05:14.727: triage child A created.
    • 03:05:44.823: A created review child B through delegate_task(mode: "async"), using claude-proxy / claude-opus-5-5.
    • 03:06:04.809: A's original run completed with its one-line escalation result.
    • 03:08:07.480: B created implementation child C through delegate_task(mode: "async"), using codex-proxy / gpt-6.1-sol.
    • 03:08:48.977: B's original run completed; C continued running.
    • 03:10:14.690: the outer delegate_task tool result was a transport error, with no task handle:
    tool call error: tool call failed for `t3-code/delegate_task`
    
    Caused by:
        timed out awaiting tools/call after 300s
    
    • 03:10:18.120: the parent concluded, Triage child failed: delegate_task tool call timed out after 300 seconds.
    • A subsequent read found C still running, with activity updated at 03:14:27.358. The transport timeout had not stopped the descendant work.

    These facts were read from the durable parent, child, grandchild, and implementation-thread timelines. No isolated reproduction was run, and no changes were made to those threads or schedules.

    Source check

    The local upstream checkout inspected was clean main at 3143335fc3a568cbbb5174764961889272135cb9. That checkout is not asserted to be the exact installed nightly commit.

    • OrchestratorMcpService.ts:99 still sets a 10-minute default / 60-minute maximum wait budget.
    • waitForTask waits for terminal task status, not merely the original child run ending.
    • delegatedTaskProgress deliberately reports waiting_for_children while nested tasks or their completion deliveries remain outstanding. Thus A's short original run does not imply the outer wait can finish.
    • CodexAdapterV2.ts:1338 supplies the MCP endpoint and authorization header without an explicit tool timeout override.
    • delegateTask returns the handle only after the blocking wait, and changes completionWake from settled_only to always on the server timeout path.

    This adds a Codex-client occurrence to the original Pi/undici report; the exact Codex transport error establishes the 300-second client limit here, without attributing its implementation to undici. The bug is the client/server wait-budget mismatch and lost recovery envelope. Waiting for nested work itself is intentional. No duplicate dispatch or permanent result loss was established in this occurrence.

    Reproduction recipe and workaround

    Use a Codex parent to delegate A with mode: "wait", timeoutMs: 600000; have A start B asynchronously and return, and have B leave a descendant running beyond five minutes. In this observed run, the parent lost its outer call at 300 seconds while descendant work continued.

    For these scheduled checks, using outer mode: "async" would return the handle immediately and avoid the long blocking call. This is a source-supported workaround, not a change made to the user's schedules or separately tested here. The graceful wait path needs to return taskId and waitTimedOut before the client deadline, or otherwise reconcile the transport budget with the advertised server wait.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions