test(server): record Claude background task and subagent lifecycles live - #13517
Conversation
Unit tests for Claude background work build frames the CLI never sends: results without terminal_reason, task_started without tool_use_id, SendMessage with agent_id instead of to. This records the flows live instead and replays them through the whole orchestrator. - claude_background_task_wake: a background Bash wake becomes exactly one continuation run; a later user message runs as its own turn. - claude_background_subagent_lifecycle: two background subagents; one wakes the root, one is stopped with TaskStop, then the first is resumed with SendMessage and wakes the root again. - claude_background_task_interrupt: interrupting a turn that owns a background Bash task clears the roster and starts no continuation. The recorder can now wait for several wake turns per prompt (told apart by their task-notification origin, since a wake queued during a turn can run before the next prompt's turn), keeps recording wakes that start right after, and can interrupt after the Nth root tool use. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oster A background subagent's own Bash steps arrive as local_bash task_started frames with is_backgrounded false. The adapter's incremental roster fallback admitted every local_bash task, so after the root turn settled those steps were marked wake-eligible: their notifications opened the continuation early and replaced its detail, so the continuation run for the claude_background_subagent_after_root recording said "Sleep 3 seconds then echo SUB_DONE_1" instead of the subagent's summary SUB_FINAL_REPORT. A foreground task blocks its tool call, so it is not background work. The fallback now skips is_backgrounded false; a task moved to the background later still reaches the roster through background_tasks_changed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| if (sdkMessageHasRootToolUse(replayMessage)) { | ||
| toolUses += 1; | ||
| if (toolUses >= input.toolUseCount) { | ||
| return; | ||
| } | ||
| } |
There was a problem hiding this comment.
🟡 Medium Adapters/ClaudeAdapterV2.testkit.ts:1319
When a root assistant frame contains two tool_use parts, recordMessagesUntilToolUse increments toolUses only once, so interruptAfterToolUses: 2 continues past the requested second tool use and may fail when no later assistant frame exists. Count the root tool_use content parts rather than assistant frames.
| if (sdkMessageHasRootToolUse(replayMessage)) { | |
| toolUses += 1; | |
| if (toolUses >= input.toolUseCount) { | |
| return; | |
| } | |
| } | |
| if (sdkMessageHasRootToolUse(replayMessage)) { | |
| toolUses += replayMessage.type === "assistant" | |
| ? replayMessage.message.content.filter((part) => part.type === "tool_use").length | |
| : 0; | |
| if (toolUses >= input.toolUseCount) { | |
| return; | |
| } | |
| } |
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.testkit.ts around lines 1319-1324:
When a root assistant frame contains two `tool_use` parts, `recordMessagesUntilToolUse` increments `toolUses` only once, so `interruptAfterToolUses: 2` continues past the requested second tool use and may fail when no later assistant frame exists. Count the root `tool_use` content parts rather than assistant frames.
There was a problem hiding this comment.
🟡 Medium
queryMode: "interrupt_restart" ignores interruptAfterToolUses, so with interruptAfter: "tool_use" the replay stops after the default one tool use even though its metadata reports the requested count. Forward interruptAfterToolUses to recordClaudeInterruptRestartQuery as well.
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.testkit.ts around line 2653:
`queryMode: "interrupt_restart"` ignores `interruptAfterToolUses`, so with `interruptAfter: "tool_use"` the replay stops after the default one tool use even though its metadata reports the requested count. Forward `interruptAfterToolUses` to `recordClaudeInterruptRestartQuery` as well.
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
ApprovabilityVerdict: Would Approve Macroscope's review found this PR approvable — This PR is mainly a replay-fixture/test-harness expansion with a narrow Claude adapter fix that prevents foreground subagent commands from being treated as background work. Two unresolved medium-severity recorder correctness findings remain, covering multi-tool counting and interrupt-restart propagation, and those findings block approval until resolved. Not approved because:
Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more. |
…er (#13519) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The background subagent lifecycle recording is the first fixture with TaskStop and SendMessage tool uses, and the native tool table did not list them, so the fixture tool-classification check failed. Both are command-like dynamic tools. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| // A foreground task blocks its tool call (a subagent's own Bash | ||
| // steps included), so it is not background work; one moved to | ||
| // the background later arrives in background_tasks_changed. | ||
| if (!isClaudeNonSubagentTask(message) || message.is_backgrounded === false) { |
There was a problem hiding this comment.
This changes the background-task roster behavior for task_started frames with is_backgrounded: false, but the existing roster test only exercises is_backgrounded: true. Could you add a focused adapter test that sends a non-subagent foreground task_started frame and asserts it does not populate pendingBackgroundTasks (while a background frame still does)?
Posted via Macroscope — Effect Service Conventions
…ive (#13517) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The unit tests for Claude background work (background Bash tasks and background subagents) build frames by hand, and the real CLI never sends some of those shapes: results without
terminal_reason,task_startedwithouttool_use_id, SendMessage withagent_idinstead ofto. This PR records the same flows live and replays them through the whole orchestrator. The recordings also exposed an adapter bug, fixed in its own commit.New live fixtures (claude-sonnet-4-6, Claude Code 2.1.281)
claude_background_task_wake: a backgroundsleep 8wakes Claude after the root turn settled. The wake becomes exactly one continuation run, and it carries the notification summary. A later user message runs as its own turn, with no wake text in it. The roster lists the task while it runs and clears on completion, and no subagent is projected.claude_background_subagent_lifecycle: two background subagents. Agent A (modelhaiku) finishes and wakes the root. Agent B finishes, then is stopped withTaskStop. Agent A is then resumed withSendMessageand wakes the root again. That is 7 runs in total; each wake is its own continuation. Agent A endscompletedwith the resumed result, the observed Haiku model, one child thread holding both answers, and attribution to the run that resumed it. Agent B endscancelled, and no subagent ever reaches the roster.claude_background_task_interrupt: a backgroundsleep 30, then a foreground command, interrupted after the second tool use. The run is interrupted, the roster clears, and no continuation follows.What the recordings settled
task_started.SendMessageto a finished subagent re-emitstask_startedwith the same task id. It carries the SendMessage call'stool_use_id, and so does the resumed run'stask_notification. The resumed child frames keep the original Agenttool_use_idasparent_tool_use_id. The adapter already assumed this; the fixture now asserts it on the recorded frames. SendMessage input is{to, summary, message}.origin.kind === "task-notification"rather than by arrival order. The lifecycle fixture waits for each continuation before sending the next message, as the recording did. When a user message is queued while a wake is pending, the wake's reply lands on the user's run. That is a real ordering hazard, left for a follow-up and not covered here.local_bashtoo. They arrive withowned_by_subagent: true, is_backgrounded: false.Fix: a subagent's foreground Bash steps stay off the roster
The adapter's incremental roster fallback admitted every
local_bashtask_started. That included a background subagent's own foreground Bash steps. After the root turn settled, those steps became wake-eligible. Their notifications opened the continuation early and overwrote its detail. In the existingclaude_background_subagent_after_rootrecording, the continuation saidSleep 3 seconds then echo SUB_DONE_1instead of the subagent's summarySUB_FINAL_REPORT. The fallback now skipsis_backgrounded: false; a task moved to the background later still reaches the roster throughbackground_tasks_changed. The after-root fixture now asserts the summary and an empty roster, and it fails without the fix.Recorder
backgroundWakeCountsreplacesawaitBackgroundWake. It sets, per prompt, how many wake turns to wait for, and it also records a wake that starts within 5s after them, so the next prompt is never offered while one is queued. Wake results are labelledresult:background-wake:N.interruptAfterToolUsesinterrupts after the Nth root tool use.Verification
vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts -t claude: 24 passed.vp test runonClaudeAdapterV2.test.ts,ClaudeAdapterV2.testkit.test.ts,OrchestratorReplayFixtures.contract.test.ts, andscripts/claudeReplayRecordingConfig.test.ts: 148 passed.claude_background_task_interrupttask_startedreopen:claude_background_subagent_lifecycleclaude_background_subagent_lifecycleclaude_background_subagent_after_rootclaude_background_task_wakeclaude_background_subagent_after_rootvp exec tsc --noEmit -p .inapps/server: 0error TS, 0warning TS.vp run knip:check: passes.vp linton the touched files: two warnings, both already present on the base, formakeClaudeAgentSdkReplayQueryRunnerLayerandlayer.Model: Claude Opus 5.5 (Claude Code)
🤖 Generated with Claude Code