test(server): drop Claude background unit tests the live fixtures cover - #13519
Conversation
Eight hand-built-frame tests in ClaudeAdapterV2.test.ts restated flows the live recordings now replay through the whole orchestrator. For each, the adapter rule it guards was reverted and a fixture failed: - authoritative background_tasks_changed roster: claude_background_task_after_root - interrupted turn clears the roster: claude_background_task_interrupt - buffered wake drains into a continuation, and a task-notification wake result settles it: claude_background_task_wake - background subagent completing after root settle, attributed to its launch run: claude_background_subagent_after_root - SendMessage resume re-opens the subagent and hydrates its second result: claude_background_subagent_lifecycle - empty roster level before the notification keeps wake eligibility: claude_background_task_wake - mixed snapshot admits only local_bash: claude_background_subagent_after_root and claude_background_subagent_lifecycle The fixtures gain the assertions those tests carried (continuation detail, subagent run and node attribution). Two sub-checks go without a fixture because the CLI has not been seen producing them: a task_progress after a subagent's notification (none in 5,565 logged progress frames) and a duplicate local_bash notification (logged duplicates are subagents only). The remaining background tests cover orderings the CLI emits nondeterministically, failures that cannot be produced on demand, and the idle-release probe. Their frames now match what the CLI sends: terminal_reason on results, tool_use_id and is_backgrounded on task_started, tool_use_id on task_notification, message.id on assistant frames, one content block per assistant frame, SendMessage input with `to`, the resumed notification under the SendMessage tool_use_id, and background Bash ids and texts taken from the recording. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ApprovabilityVerdict: Approved at Macroscope's review found this PR approvable — This PR is confined to test-only unit and replay-fixture assertions: it removes duplicated synthetic-frame coverage and strengthens live-fixture checks without touching production code, defaults, or deployment configuration. Its runtime blast radius is limited to test execution. Notes:
You can add or adjust custom eligibility rules. Learn more. |
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
The resume test's ACK text dropped the short agent id and the pin object the recorded frame carries, and omitted tool_use_result. Build both from one object shaped like the claude_background_subagent_lifecycle recording. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Dismissing prior approval to re-evaluate 419592a
Stacked on #13517. Now that the live fixtures cover Claude background work, the hand-built-frame unit tests that restated the same flows can go. The tests that stay should use frames the CLI actually sends.
Deleted: 8 tests, about 1,030 lines of
ClaudeAdapterV2.test.tsFor each test, the adapter rule it guards was reverted and the named fixture failed:
background_tasks_changedrosterclaude_background_task_after_rootclaude_background_task_interruptclaude_background_task_wakeclaude_background_task_wakeclaude_background_subagent_after_rootclaude_background_subagent_lifecycleclaude_background_task_wakebackground_tasks_changedsnapshotclaude_background_subagent_after_root,claude_background_subagent_lifecycleThe fixtures pick up the assertions those tests made: the continuation detail, and the subagent's run and node attribution. Two sub-checks go without a fixture, because the CLI has not been seen producing them:
task_progressafter a subagent's notification (none among 5,565 logged progress frames)local_bashnotification (the logged duplicates are all subagents)Kept, with realistic frames
The remaining background tests cover three kinds of behavior:
task_started, a resumetask_startedracing past settle, interrupt races)Their frames now match what the CLI sends:
terminal_reasonon results, except a zero-turn task-notification result, which omits it as seen livetool_use_idandis_backgroundedontask_startedtool_use_idontask_notificationmessage.idon assistant frames, one content block per frame{to, summary, message}, the resume ACK text as recorded, and the resumed notification under the SendMessagetool_use_idclaude_background_task_wakerecordingThe held-frame test still sends a
task_startedwithouttool_use_id. The SDK types that field as optional, but no recorded or logged frame has omitted it. The test now says so.Verification
vp test runonClaudeAdapterV2.test.ts,ClaudeAdapterV2.testkit.test.ts,OrchestratorReplayFixtures.contract.test.ts, andscripts/claudeReplayRecordingConfig.test.ts: 140 passed.vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts -t claude: 24 passed.vp exec tsc --noEmit -p .inapps/server: 0error TS, 0warning TS.vp run knip:check: passes.vp linton the touched files: clean.Model: Claude Opus 5.5 (Claude Code)
🤖 Generated with Claude Code