Skip to content

test(server): record Claude background task and subagent lifecycles live - #13517

Merged
juliusmarminge merged 4 commits into
t3code/codex-turn-mappingfrom
v2/claude-background
Sep 25, 2026
Merged

juliusmarminge merged 4 commits into
t3code/codex-turn-mappingfrom
v2/claude-background

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

The unit tests for Claude background work (background Bash tasks and background subagents) build frames by hand, and the real CLI never sends some of those shapes: results without terminal_reason, task_started without tool_use_id, SendMessage with agent_id instead of to. This PR records the same flows live and replays them through the whole orchestrator. The recordings also exposed an adapter bug, fixed in its own commit.

New live fixtures (claude-sonnet-4-6, Claude Code 2.1.281)

  • claude_background_task_wake: a background sleep 8 wakes Claude after the root turn settled. The wake becomes exactly one continuation run, and it carries the notification summary. A later user message runs as its own turn, with no wake text in it. The roster lists the task while it runs and clears on completion, and no subagent is projected.
  • claude_background_subagent_lifecycle: two background subagents. Agent A (model haiku) finishes and wakes the root. Agent B finishes, then is stopped with TaskStop. Agent A is then resumed with SendMessage and wakes the root again. That is 7 runs in total; each wake is its own continuation. Agent A ends completed with the resumed result, the observed Haiku model, one child thread holding both answers, and attribution to the run that resumed it. Agent B ends cancelled, and no subagent ever reaches the roster.
  • claude_background_task_interrupt: a background sleep 30, then a foreground command, interrupted after the second tool use. The run is interrupted, the roster clears, and no continuation follows.

What the recordings settled

  • Resume re-emits task_started. SendMessage to a finished subagent re-emits task_started with the same task id. It carries the SendMessage call's tool_use_id, and so does the resumed run's task_notification. The resumed child frames keep the original Agent tool_use_id as parent_tool_use_id. The adapter already assumed this; the fixture now asserts it on the recorded frames. SendMessage input is {to, summary, message}.
  • Wake turns can run before a queued prompt. A notification that lands during an active turn queues a wake turn, and the CLI runs it before the next user prompt's turn. The recorder therefore tells results apart by origin.kind === "task-notification" rather than by arrival order. The lifecycle fixture waits for each continuation before sending the next message, as the recording did. When a user message is queued while a wake is pending, the wake's reply lands on the user's run. That is a real ordering hazard, left for a follow-up and not covered here.
  • A subagent's foreground Bash steps are local_bash too. They arrive with owned_by_subagent: true, is_backgrounded: false.

Fix: a subagent's foreground Bash steps stay off the roster

The adapter's incremental roster fallback admitted every local_bash task_started. That included a background subagent's own foreground Bash steps. After the root turn settled, those steps became wake-eligible. Their notifications opened the continuation early and overwrote its detail. In the existing claude_background_subagent_after_root recording, the continuation said Sleep 3 seconds then echo SUB_DONE_1 instead of the subagent's summary SUB_FINAL_REPORT. The fallback now skips is_backgrounded: false; a task moved to the background later still reaches the roster through background_tasks_changed. The after-root fixture now asserts the summary and an empty roster, and it fails without the fix.

Recorder

backgroundWakeCounts replaces awaitBackgroundWake. It sets, per prompt, how many wake turns to wait for, and it also records a wake that starts within 5s after them, so the next prompt is never offered while one is queued. Wake results are labelled result:background-wake:N. interruptAfterToolUses interrupts after the Nth root tool use.

Verification

  • vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts -t claude: 24 passed.
  • vp test run on ClaudeAdapterV2.test.ts, ClaudeAdapterV2.testkit.test.ts, OrchestratorReplayFixtures.contract.test.ts, and scripts/claudeReplayRecordingConfig.test.ts: 148 passed.
  • Revert checks. For each, one adapter rule was reverted and the named fixture failed:
    • roster clear on interrupted turns: claude_background_task_interrupt
    • task_started reopen: claude_background_subagent_lifecycle
    • async-launch ACK and SendMessage ACK handling: claude_background_subagent_lifecycle
    • single continuation offer: claude_background_subagent_after_root
    • wake drain: three fixtures
    • an empty roster level clearing wake eligibility: claude_background_task_wake
    • the new foreground fix: claude_background_subagent_after_root
  • vp exec tsc --noEmit -p . in apps/server: 0 error TS, 0 warning TS.
  • vp run knip:check: passes.
  • vp lint on the touched files: two warnings, both already present on the base, for makeClaudeAgentSdkReplayQueryRunnerLayer and layer.
  • Not run: repo-wide checks.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

juliusmarminge and others added 2 commits September 24, 2026 16:21
Unit tests for Claude background work build frames the CLI never sends:
results without terminal_reason, task_started without tool_use_id,
SendMessage with agent_id instead of to. This records the flows live
instead and replays them through the whole orchestrator.

- claude_background_task_wake: a background Bash wake becomes exactly one
  continuation run; a later user message runs as its own turn.
- claude_background_subagent_lifecycle: two background subagents; one wakes
  the root, one is stopped with TaskStop, then the first is resumed with
  SendMessage and wakes the root again.
- claude_background_task_interrupt: interrupting a turn that owns a
  background Bash task clears the roster and starts no continuation.

The recorder can now wait for several wake turns per prompt (told apart by
their task-notification origin, since a wake queued during a turn can run
before the next prompt's turn), keeps recording wakes that start right
after, and can interrupt after the Nth root tool use.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oster

A background subagent's own Bash steps arrive as local_bash task_started
frames with is_backgrounded false. The adapter's incremental roster
fallback admitted every local_bash task, so after the root turn settled
those steps were marked wake-eligible: their notifications opened the
continuation early and replaced its detail, so the continuation run for
the claude_background_subagent_after_root recording said "Sleep 3
seconds then echo SUB_DONE_1" instead of the subagent's summary
SUB_FINAL_REPORT.

A foreground task blocks its tool call, so it is not background work.
The fallback now skips is_backgrounded false; a task moved to the
background later still reaches the roster through
background_tasks_changed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added the vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. label Sep 24, 2026
@github-actions github-actions Bot added the size:XXL 1,000+ changed lines (additions + deletions). label Sep 24, 2026
Comment on lines +1319 to 1324
if (sdkMessageHasRootToolUse(replayMessage)) {
toolUses += 1;
if (toolUses >= input.toolUseCount) {
return;
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium Adapters/ClaudeAdapterV2.testkit.ts:1319

When a root assistant frame contains two tool_use parts, recordMessagesUntilToolUse increments toolUses only once, so interruptAfterToolUses: 2 continues past the requested second tool use and may fail when no later assistant frame exists. Count the root tool_use content parts rather than assistant frames.

Suggested change
if (sdkMessageHasRootToolUse(replayMessage)) {
toolUses += 1;
if (toolUses >= input.toolUseCount) {
return;
}
}
if (sdkMessageHasRootToolUse(replayMessage)) {
toolUses += replayMessage.type === "assistant"
? replayMessage.message.content.filter((part) => part.type === "tool_use").length
: 0;
if (toolUses >= input.toolUseCount) {
return;
}
}
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.testkit.ts around lines 1319-1324:

When a root assistant frame contains two `tool_use` parts, `recordMessagesUntilToolUse` increments `toolUses` only once, so `interruptAfterToolUses: 2` continues past the requested second tool use and may fail when no later assistant frame exists. Count the root `tool_use` content parts rather than assistant frames.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium

queryMode: "interrupt_restart" ignores interruptAfterToolUses, so with interruptAfter: "tool_use" the replay stops after the default one tool use even though its metadata reports the requested count. Forward interruptAfterToolUses to recordClaudeInterruptRestartQuery as well.

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.testkit.ts around line 2653:

`queryMode: "interrupt_restart"` ignores `interruptAfterToolUses`, so with `interruptAfter: "tool_use"` the replay stops after the default one tool use even though its metadata reports the requested count. Forward `interruptAfterToolUses` to `recordClaudeInterruptRestartQuery` as well.

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.1 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 1 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.8 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: e6a4348 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@macroscopeapp

macroscopeapp Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Would Approve

Macroscope's review found this PR approvable — This PR is mainly a replay-fixture/test-harness expansion with a narrow Claude adapter fix that prevents foreground subagent commands from being treated as background work. Two unresolved medium-severity recorder correctness findings remain, covering multi-tool counting and interrupt-restart propagation, and those findings block approval until resolved.

Not approved because:

  • 2 blocking correctness issues found at or above your repo's Minimum Blocking Severity

Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more.

juliusmarminge and others added 2 commits September 24, 2026 17:47
…er (#13519)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The background subagent lifecycle recording is the first fixture with
TaskStop and SendMessage tool uses, and the native tool table did not
list them, so the fixture tool-classification check failed. Both are
command-like dynamic tools.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
// A foreground task blocks its tool call (a subagent's own Bash
// steps included), so it is not background work; one moved to
// the background later arrives in background_tasks_changed.
if (!isClaudeNonSubagentTask(message) || message.is_backgrounded === false) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This changes the background-task roster behavior for task_started frames with is_backgrounded: false, but the existing roster test only exercises is_backgrounded: true. Could you add a focused adapter test that sends a non-subagent foreground task_started frame and asserts it does not populate pendingBackgroundTasks (while a background frame still does)?

Posted via Macroscope — Effect Service Conventions

@juliusmarminge
juliusmarminge merged commit 22da2ee into t3code/codex-turn-mapping Sep 25, 2026
24 checks passed
@juliusmarminge
juliusmarminge deleted the v2/claude-background branch September 25, 2026 01:06
juliusmarminge added a commit that referenced this pull request Sep 25, 2026
…ive (#13517)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL 1,000+ changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant