Skip to content

fix(server): keep a Claude wake reply off the user's queued message - #13535

Merged
juliusmarminge merged 3 commits into
t3code/codex-turn-mappingfrom
v2/claude-queued-wake-order
Sep 25, 2026
Merged

juliusmarminge merged 3 commits into
t3code/codex-turn-mappingfrom
v2/claude-queued-wake-order

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

When a background task or subagent finishes during a turn, Claude queues a wake turn and runs it before the next prompt's turn. The adapter took whichever turn came next as the answer to the prompt. So a message sent right after such a turn showed the wake reply, and its real answer landed in the following continuation run.

Before / after (same live recording, replayed)

Message Before After
"Use the SendMessage tool to send Agent A…" (resume) "Agent B's kill is confirmed by the notification. No further action needed — the task is complete." RESUMED
continuation run RESUMED "Agent B's kill is confirmed by the notification…"

What the live CLI shows (Claude Code 2.1.281)

  • The SDK's SDKUserMessage.uuid is echoed back as user_message_uuid (plus user_message_uuids). It appears on the first root message_start of the turn that answers the prompt, on its first assistant frame, and on its result.
  • Wake turns carry no echo at all.
  • The CLI also sends an undeclared command_lifecycle frame (queued / started / completed, keyed by command_uuid) for each prompt that carries a uuid.
  • In the new recording, prompt.offer:3 is followed by a wake turn (result origin: task-notification, no echo). Only then comes the prompt's turn, whose first message_start echoes its uuid.
  • The provider logs agree. On 2.1.280/2.1.281, 76 of 78 user turns that sent a uuid stamp it on their first message_start, and none of 89 wake turns carry one. 2.1.236 echoes only on the result.

Fix

  • Prompt uuid. Each prompt now carries a uuid derived from its run attempt (claudePromptUuid, SHA-256 shaped as a v4 uuid), so replays stay deterministic.
  • Learning the echo mode. Each CLI process learns from its first prompt turn whether the uuid comes back early. That first turn can't have a wake queued ahead of it.
  • Hold only when the echo comes early. Once early echo is known, the output of a turn that arrives before a prompt's echo is held.
  • The held turn keeps everything it produced. Held with it: its root frames; the subagent and task-lifecycle frames tied to its tool uses (kept in order); and anything its permission callbacks produce.
    • A callback that needs no answer defers to the held tool_use frame. That frame starts the tool, and an ExitPlanMode plan is projected when the frame is handled, in whichever run it lands in. This holds in every runtime mode: ExitPlanMode answers at once, so it never takes the approval path below. Its plan is deferred only while its tool_use frame is still held and is published immediately otherwise, so a plan Claude is told was captured is never dropped.
    • A callback that needs the user's answer (an approval, AskUserQuestion) cannot wait for an echo that only comes after it is answered. It releases the held output to the pending prompt turn and asks there, which is where this output went before the hold existed. So the SDK is never left waiting.
  • Where held output goes:
    • The echo releases it to the prompt's turn.
    • A task-notification-origin result with at least one turn sends it to the wake buffer, which becomes a continuation run as before.
    • A zero-turn task-notification result is lifecycle debris: it is dropped as before, and held output stays held for its real owner.
    • Any other result settles the prompt's turn with the held output. That includes an interrupt or a failure, and turns of another origin that can run ahead (peer, channel, coordinator, auto-continuation). Their attribution is unchanged from before; only their streaming is delayed until their result.
  • No delay on the common path. The prompt's own turn echoes on its first frame, so its streaming is never held; only a turn running ahead of it is held.
  • CLIs that never echo early (2.1.236-style) are never held and behave exactly as before, including the misattribution itself. A fixture pins that.
  • Stream end. If the stream ends while output is held, the held output goes to the turn before it is finalized. A replaced query's exit leaves the live turn alone.

Replay maps recorded prompt uuids to the replayed ones and accepts command_lifecycle frames; older recordings without uuids still match. The recorder now sends uuids too, and can offer the next prompt without waiting out queued wakes.

Tests

  • claude_background_wake_before_queued_prompt (fixture), re-recorded live on 2.1.281 with uuids:
    • it first asserts the ordering in the recorded frames;
    • then each of the four user messages gets exactly its own reply, and both wake replies land in continuation runs.
  • claude_background_wake_before_queued_prompt_no_echo (fixture): the earlier recording made without uuids. It pins that a CLI without an echo streams as before.
  • Adapter tests in ClaudeAdapterV2.test.ts:
    • A queued wake turn's ExitPlanMode callback (the review's case), in both full-access and approval-required mode. Driven through continuation drain: the tool (every update), its proposed plan, and the plan artifact all land in the continuation run. The tool is not failed by the user's run, the user's run shows only its own reply, and the plan Claude is told was captured is projected.
    • A subagent a queued wake turn launches. Its lifecycle and child frames stay with the wake turn's continuation.
    • An approval a held wake turn raises. It is answered without waiting for the echo.
    • A zero-turn task-notification result before the echo. No continuation is opened.
    • Held output released when the stream ends.

Verification

  • vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts -t claude: 26 passed.
  • vp test run on ClaudeAdapterV2.test.ts, ClaudeReplayFixtures.integration.test.ts, and OrchestratorReplayFixtures.contract.test.ts: 122 passed.
  • Fail-first checks:
    • The fixture fails with the gate disabled: the resume message gets the wake reply instead of RESUMED.
    • On the first commit of this PR, the ExitPlanMode test, the wake-subagent test, and the zero-turn test all fail.
    • On b4b4009d8f, the approval-required ExitPlanMode case fails (its plan is never published); the full-access case passes.
    • The stream-end test fails without the release.
  • vp exec tsc --noEmit -p . in apps/server: 0 error TS, 0 warning TS.
  • vp run knip:check: passes.
  • vp lint on the touched files: only the one warning already on the base.
  • Not run: repo-wide checks.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code

@macroscopeapp

macroscopeapp Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting).

This review would cost an estimated $15.76, which exceeds your per-review limit of $15.00.

The top 3 files driving up this estimate:

File Diff Size Estimate
apps/server/src/orchestration-v2/testkit/fixtures/claude_background_wake_before_queued_prompt/claude_transcript.ndjson 152.15KB $7.61
apps/server/src/orchestration-v2/testkit/fixtures/claude_background_wake_before_queued_prompt_no_echo/claude_transcript.ndjson 124.82KB $6.24
apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.ts 18.73KB $0.94

Tip

To get this pull request reviewed, you can:

  1. Comment @macroscope-app on this PR to request a manual review (monthly spend limits still apply).
  2. Exclude the file(s) above from review by adding a pattern to your .macroscope/ignore.md — note that creating this file replaces Macroscope's built-in default ignores rather than extending them.
  3. Raise your cost limit in your workspace billing settings.

Turn off this reminder going forward

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Sep 25, 2026
@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.8 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: 40a8a58 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@macroscopeapp

macroscopeapp Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR makes substantial production changes to Claude turn attribution, buffering, continuation routing, tool callbacks, and subagent lifecycle handling, backed by extensive new fixtures and tests. The asynchronous state-machine behavior and broad runtime impact warrant human review.

Not approved because:

  • Per-review cost limit exceeded (workspace setting). Approvability relies on correctness review in order to determine eligibility

Review your spending limits in Billing settings, or comment @macroscope-app review this PR to bypass the limit and review now. You can add or adjust custom eligibility rules. Learn more.

juliusmarminge and others added 2 commits September 24, 2026 19:48
When a background task or subagent finishes during a turn, Claude queues
a wake turn and runs it before the next prompt's turn. The adapter took
whichever turn came next as the answer to the prompt, so a message sent
right after such a turn showed the wake reply ("Agent B's kill is
confirmed…") and its real answer landed in the next continuation run.

The adapter now sends each prompt with a uuid derived from its run
attempt. Claude Code echoes it as user_message_uuid on the first frame of
the turn that answers it; wake turns carry none. Once a CLI process has
echoed on a turn's first frame, root output arriving before a prompt's
echo is held: the echo releases it to the prompt's turn, and a
task-notification-origin result sends it to the wake buffer and so to a
continuation run. The prompt's own turn is never delayed, because its
echo is on its first frame. A CLI that echoes only on the result, or not
at all, is never held and behaves as before. Held output is released to
the turn if the stream ends first.

Replay maps recorded prompt uuids to the replayed ones and accepts the
command_lifecycle frames Claude Code 2.1.281 sends for uuid-carrying
prompts. The recorder sends uuids too and can offer the next prompt
without waiting out queued wakes.

- claude_background_wake_before_queued_prompt: re-recorded live on
  2.1.281 with uuids; each prompt gets its own reply and both wake
  replies land in continuation runs. Fails without the gate.
- claude_background_wake_before_queued_prompt_no_echo: the earlier
  recording without uuids, pinning that a CLI without an echo streams
  as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
While a prompt's echo is pending, a queued wake turn's root output is held
and later goes to a continuation run. But its permission callbacks and
the frames of subagents it launched were still attached to the pending
prompt turn: an ExitPlanMode in a wake turn started its tool and plan in
the user's run, the tool was failed when that run settled, and the drain
moved the tool while the plan stayed behind.

- A permission callback that needs no answer defers to the held tool_use
  frame: the frame starts the tool, and an ExitPlanMode plan projects when
  that frame is handled, in whichever run it lands in.
- A callback that needs the user's answer (an approval, AskUserQuestion)
  cannot wait for an echo that only comes after it is answered, so it
  releases the held output to the prompt turn and asks there, as before
  output was held.
- Subagent and task lifecycle frames tied to a held tool_use are held with
  it, in order.
- A zero-turn task-notification result before the echo is lifecycle debris
  again: it is dropped as before instead of opening an empty
  continuation.
- A replaced query's stream end no longer releases the live turn's held
  output.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@juliusmarminge
juliusmarminge force-pushed the v2/claude-queued-wake-order branch from 7058060 to b4b4009 Compare September 25, 2026 02:52
In approval-required mode a held ExitPlanMode callback went through the
interactive-approval fallback, which released the held frames, and then
parked the plan for a tool_use frame that had already been handled, so
the plan was never published while Claude was told it was captured.

ExitPlanMode answers at once and needs no user input, so it no longer
takes that fallback. Its plan is deferred only while its tool_use frame is
still held, and published immediately otherwise.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@juliusmarminge
juliusmarminge merged commit 1ee0577 into t3code/codex-turn-mapping Sep 25, 2026
24 of 25 checks passed
@juliusmarminge
juliusmarminge deleted the v2/claude-queued-wake-order branch September 25, 2026 03:34
juliusmarminge added a commit that referenced this pull request Sep 25, 2026
…13535)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL 1,000+ changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant