Skip to content

fix(server): Claude subagent threads show their thinking - #13626

Merged
juliusmarminge merged 1 commit into
t3code/codex-turn-mappingfrom
v2/claude-subagent-thinking
Sep 25, 2026
Merged

juliusmarminge merged 1 commit into
t3code/codex-turn-mappingfrom
v2/claude-subagent-thinking

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

A native Claude subagent's child thread showed its tool calls and final reply but none of the thinking in between, even though Claude sends it.

Why. Claude never streams subagent output (0 stream_event frames with a parent_tool_use_id across the provider logs we checked). Instead, each subagent thinking block arrives whole in an assistant snapshot tagged with the subagent's parent_tool_use_id. ClaudeAdapterV2 dropped those blocks in two places:

  • the reasoning path skips frames that have a parent_tool_use_id, which is correct for the parent thread;
  • subagent text routing uses textFromClaudeContent, which keeps only text parts.

What changed

Subagent thinking blocks now become completed reasoning items titled "Thinking" in the subagent's child thread. They go through the same resolveSubagentByToolUseId path and child ordinal counter as subagent text, so they sit before the reply they led to. Empty or whitespace-only blocks are skipped, and frames whose owner isn't registered yet are still held by the existing pending-frame buffer.

Verification

  • The existing live recording claude_background_subagent_lifecycle already has subagent thinking (3 blocks across 2 subagents, including a SendMessage resume), so I didn't record anything new. Its output.ts now asserts:
    • Agent A's child has 2 thinking items that reference "A_FIRST" and "A_SECOND", ordered thinking → A_FIRST → thinking → A_SECOND across the resume, each completed, not streaming, with runId: null.
    • Agent B's child has the thinking about its B_DONE command.
    • None of that subagent thinking appears in the parent thread.
  • With the adapter change reverted, the fixture fails ("expected [] to have a length of 2").
  • cd apps/server && vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts src/orchestration-v2/testkit/ClaudeReplayFixtures.integration.test.ts src/orchestration-v2/Adapters/ClaudeAdapterV2.test.ts: 207 passed. The other Claude subagent recordings (claude_background_subagent_after_root, claude_background_wake_before_queued_prompt*, subagent) replay unchanged.
  • vp exec tsc --noEmit -p . (server): no errors. vp lint: no new findings (the unused layer warning is already on the base). vp run knip:check: clean.

Deliberately left out

  • task_progress usage counts (tool uses, tokens). The V2 OrchestrationV2Subagent contract has no usage field, and projectedSubagentsToRuntime hard-codes usage: null. Showing counts would need a contract change plus a client surface that renders them; no current surface does for V2. That's a separate concern.
  • Removing the "Subagent progress" reasoning item. With #13614 the child shows a live tool row and working state, but that item is still the only in-thread place the latest task_progress description appears, and the subagent fixture asserts it. Whether to drop it is a product call after fix(clients): native subagent threads show when they are working #13614 lands.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

Claude sends each subagent thinking block whole, in an assistant snapshot
tagged with the subagent's parent_tool_use_id (it never streams subagent
output). The adapter dropped those blocks: the reasoning path skips
parent-tool frames, and subagent text routing keeps only text parts. A
native subagent's thread showed its tool calls and final reply with none
of the reasoning in between.

Route subagent thinking blocks into the child thread as completed
"Thinking" reasoning items, ordered with the child's other items, the
same way subagent text already reaches it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 25, 2026
: yield* resolveSubagentByToolUseId(context, assistantParentToolUseId);
if (subagent !== undefined) {
const now = yield* DateTime.now;
for (const [index, text] of thinking.entries()) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium Adapters/ClaudeAdapterV2.ts:5378

Duplicate assistant snapshots with the same message.uuid emit the same derived reasoning item ID with a newly incremented nextChildItemOrdinal, so later text or tool items can sort before the thinking they followed. Deduplicate subagent reasoning snapshots before this loop, mirroring the root path's context.reasoning.snapshots handling.

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.ts around line 5378:

Duplicate assistant snapshots with the same `message.uuid` emit the same derived reasoning item ID with a newly incremented `nextChildItemOrdinal`, so later text or tool items can sort before the thinking they followed. Deduplicate subagent reasoning snapshots before this loop, mirroring the root path's `context.reasoning.snapshots` handling.

@macroscopeapp

macroscopeapp Bot commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Would Approve

Macroscope's review found this PR approvable — This is a localized Claude orchestration bug fix that projects subagent thinking into child threads and adds focused lifecycle assertions without schema, deployment, or sensitive-path changes. A separate unresolved Medium correctness finding flags possible duplicate-snapshot ordering problems and should be handled by the correctness gate.

Not approved because:

  • 1 blocking correctness issue found at or above your repo's Minimum Blocking Severity

Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more.

@github-actions

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.7 KiB — 29.3 KiB ✅
Claude Live turn messages — 1 — 8 ✅

Baseline: unavailable · PR result: fa18ccd · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@juliusmarminge
juliusmarminge merged commit 690c57f into t3code/codex-turn-mapping Sep 25, 2026
24 of 25 checks passed
@juliusmarminge
juliusmarminge deleted the v2/claude-subagent-thinking branch September 25, 2026 19:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M 30-99 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant