fix(server): Claude subagent threads show their thinking - #13626
Conversation
Claude sends each subagent thinking block whole, in an assistant snapshot tagged with the subagent's parent_tool_use_id (it never streams subagent output). The adapter dropped those blocks: the reasoning path skips parent-tool frames, and subagent text routing keeps only text parts. A native subagent's thread showed its tool calls and final reply with none of the reasoning in between. Route subagent thinking blocks into the child thread as completed "Thinking" reasoning items, ordered with the child's other items, the same way subagent text already reaches it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| : yield* resolveSubagentByToolUseId(context, assistantParentToolUseId); | ||
| if (subagent !== undefined) { | ||
| const now = yield* DateTime.now; | ||
| for (const [index, text] of thinking.entries()) { |
There was a problem hiding this comment.
🟡 Medium Adapters/ClaudeAdapterV2.ts:5378
Duplicate assistant snapshots with the same message.uuid emit the same derived reasoning item ID with a newly incremented nextChildItemOrdinal, so later text or tool items can sort before the thinking they followed. Deduplicate subagent reasoning snapshots before this loop, mirroring the root path's context.reasoning.snapshots handling.
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.ts around line 5378:
Duplicate assistant snapshots with the same `message.uuid` emit the same derived reasoning item ID with a newly incremented `nextChildItemOrdinal`, so later text or tool items can sort before the thinking they followed. Deduplicate subagent reasoning snapshots before this loop, mirroring the root path's `context.reasoning.snapshots` handling.
ApprovabilityVerdict: Would Approve Macroscope's review found this PR approvable — This is a localized Claude orchestration bug fix that projects subagent thinking into child threads and adds focused lifecycle assertions without schema, deployment, or sensitive-path changes. A separate unresolved Medium correctness finding flags possible duplicate-snapshot ordering problems and should be handled by the correctness gate. Not approved because:
Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more. |
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
690c57f
into
t3code/codex-turn-mapping
A native Claude subagent's child thread showed its tool calls and final reply but none of the thinking in between, even though Claude sends it.
Why. Claude never streams subagent output (0
stream_eventframes with aparent_tool_use_idacross the provider logs we checked). Instead, each subagent thinking block arrives whole in anassistantsnapshot tagged with the subagent'sparent_tool_use_id.ClaudeAdapterV2dropped those blocks in two places:parent_tool_use_id, which is correct for the parent thread;textFromClaudeContent, which keeps onlytextparts.What changed
Subagent
thinkingblocks now become completedreasoningitems titled "Thinking" in the subagent's child thread. They go through the sameresolveSubagentByToolUseIdpath and child ordinal counter as subagent text, so they sit before the reply they led to. Empty or whitespace-only blocks are skipped, and frames whose owner isn't registered yet are still held by the existing pending-frame buffer.Verification
claude_background_subagent_lifecyclealready has subagent thinking (3 blocks across 2 subagents, including aSendMessageresume), so I didn't record anything new. Itsoutput.tsnow asserts:completed, not streaming, withrunId: null.B_DONEcommand.cd apps/server && vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts src/orchestration-v2/testkit/ClaudeReplayFixtures.integration.test.ts src/orchestration-v2/Adapters/ClaudeAdapterV2.test.ts: 207 passed. The other Claude subagent recordings (claude_background_subagent_after_root,claude_background_wake_before_queued_prompt*,subagent) replay unchanged.vp exec tsc --noEmit -p .(server): no errors.vp lint: no new findings (the unusedlayerwarning is already on the base).vp run knip:check: clean.Deliberately left out
task_progressusage counts (tool uses, tokens). The V2OrchestrationV2Subagentcontract has no usage field, andprojectedSubagentsToRuntimehard-codesusage: null. Showing counts would need a contract change plus a client surface that renders them; no current surface does for V2. That's a separate concern.task_progressdescription appears, and thesubagentfixture asserts it. Whether to drop it is a product call after fix(clients): native subagent threads show when they are working #13614 lands.Model: Claude Opus 5.5 (Claude Code)
🤖 Generated with Claude Code