Skip to content

fix(chat): show provider reasoning in the timeline - #8628

Closed
kgarg2468 wants to merge 19 commits into
pingdotgg:mainfrom
kgarg2468:t3code/streaming-reasoning-transcript
Closed

kgarg2468 wants to merge 19 commits into
pingdotgg:mainfrom
kgarg2468:t3code/streaming-reasoning-transcript

Conversation

@kgarg2468

@kgarg2468 kgarg2468 commented Aug 29, 2026 •

Copy link
Copy Markdown

Provider reasoning output never reached any client. Ingestion forwarded only assistant_text. It dropped reasoning_text and reasoning_summary_text, so Claude, Codex, and OpenCode all lost their reasoning.

This implements the fix discussed in discussion #8625 and addresses open issue #5542.

What changed

Reasoning now rides the existing assistant message pipeline through an optional channel: "reasoning" field on assistant messages. There is no parallel message type or second store. One burst of thinking becomes one message keyed reasoning:<threadId>:<turnId>:segment:<n>. Events without a turn ID use turnless in the turn slot.

Reasoning delivery follows the existing enableLegacyTokenStreaming setting. There is no new delivery mode. Per-thread monotonic timestamps keep live ordering and reload ordering identical.

The web client nests each reasoning message under the turn's Worked for Ns group as a collapsible row. A completed burst that took at least one second reads Thought for Ns. A shorter burst reads Thought.

Reasoning is excluded from answer semantics. It never settles a turn, lands in a checkpoint, supplies a thread title, appears in search, or appears in the minimap.

Migration 044_ProjectionThreadMessagesChannel adds a nullable channel column to projection_thread_messages. A PRAGMA table_info check guards the alteration. This makes the migration idempotent, which means running it twice has the same effect as running it once.

The Claude adapter now sends thinking: { type: "adaptive", display: "summarized" } unless thinking is explicitly disabled. Recent Claude Code versions otherwise stream redacted thinking that carries token estimates and no text. This removes the --thinking-display summarized launch-argument workaround described in #5542.

Evidence

Before: the timeline on current main has no reasoning row.

Timeline on current main, with no reasoning row

After: the timeline on this branch has an expanded Thought for 7.7s row.

Timeline on this branch, showing an expanded Thought for 7.7s row

Demo video: https://youtu.be/VuzwmRg24uM

Surfaces

  • Entry points: Only passive rendering in the chat view applies. There is no action to mirror in Settings, the command palette, or a keybinding.

  • Clients: Web renders the rows. Desktop gets them by wrapping web. Mobile filters reasoning out during derivation because it has no reasoning UI yet.

  • Providers: Claude, Codex, and OpenCode already normalize thinking deltas to reasoning_text and reasoning_summary_text. This change consumes both stream kinds. Cursor and Grok expose no reasoning stream, so there is nothing to show.

  • Contracts: The optional channel field crosses the typed message, command, event, persistence, and snapshot contracts.

  • Reverse states: Each row can be expanded and collapsed. There is no persisted one-way state.

  • Connection modes and version skew: The existing assistant message path serves local, remote/relay, and tunnel connections. Schema decoding strips the unknown channel field for older clients, so they render reasoning as ordinary assistant text instead of failing. An older server rolled back onto events carrying channel still replays and starts.

  • Docs: docs/user/reasoning.md explains the shipped behavior, docs/user/providers-claude.md covers Claude's adaptive summarized thinking, and docs/internals/glossary.md defines the channel.

Verification

Claude is verified end to end in a real client. Codex and OpenCode use the same ingestion path, and tests cover both reasoning stream kinds. Neither provider was manually driven.

Tests: focused runs of threadActivity.test.ts, CheckpointReactor.test.ts, ProjectionPipeline.test.ts, ProjectionSnapshotQuery.test.ts, ProviderCommandReactor.test.ts, ProviderRuntimeIngestion.test.ts, projector.test.ts, 044_ProjectionThreadMessagesChannel.test.ts, ClaudeAdapter.test.ts, MessagesTimeline.logic.test.ts, and threadReducer.test.ts, plus scoped typechecks for the touched packages. All passed.

Known boundary

The Claude adapter's content_block_start handler still does not register thinking blocks, so thinking deltas still carry no itemId. This design does not need that ID because it keys reasoning bursts by thread, turn, and segment. I can fix this at the source if maintainers prefer.

Scope and split option

This PR is larger than the usual one-concern ideal. It has 906 additions and 65 deletions across 32 files, with roughly half of the additions in tests. It is rebased on current main.

The change splits into three pieces: the two-line Claude thinking.display option, the server pipeline, and the web rows. I will reshape or split it on request.

Model: GPT-5.6 Sol. Harness: Codex, orchestrated from Claude Code.

Note

Show provider reasoning messages in the chat timeline

  • Adds a new channel: "reasoning" discriminator to orchestration messages, persisted via migration 44 and threaded through the projection, projector, decider, and client reducer
  • Runtime ingestion segments reasoning deltas per turn, supports streaming or buffered delivery modes, and finalizes orphaned segments across restarts and terminal boundaries; assistant createdAt timestamps are now strictly increasing per thread
  • Timeline renders reasoning messages as collapsible "Thinking..." / "Thought" disclosure rows that default to expanded while streaming and persist user toggles; empty completed reasoning entries are filtered from the timeline and minimap
  • Reasoning messages are excluded from turn settlement, checkpoint assistantMessageId binding, revert retention counting, thread title generation, and terminal assistant tracking across server and client
  • Claude adapter sends thinking: { type: "adaptive", display: "summarized" } by default; Codex turn-start params include summary: "detailed"
  • Risk: hasAssistantMessageForTurn and deriveTerminalAssistantMessageIds now skip reasoning messages — any code that assumed all assistant-role messages count toward a turn may need the same isReasoningMessage guard; migration 44 adds a nullable channel column to projection_thread_messages

Macroscope summarized 612deb2.


Note

Medium Risk
Touches orchestration ingestion, projection revert retention, checkpoints, and a DB migration; incorrect reasoning/answer separation could mis-bind turns or drop messages on revert, but behavior is heavily test-backed.

Overview
Adds first-class reasoning assistant messages via optional channel: "reasoning", wired from provider reasoning_text / reasoning_summary_text through ingestion, events, SQLite projection (migration 044), and snapshots.

Server ingestion segments reasoning per thread/turn (reasoning:<thread>:<turn|turnless>:segment:n), honors buffered vs streaming like normal assistant text, finalizes segments at answers, tools, pauses, and turn/session boundaries, and uses monotonic per-thread message stamps. Reasoning is excluded from answer semantics: turn settlement, checkpoint assistantMessageId, revert assistant fallback/counting, and thread-title context all skip it via isReasoningMessage.

Clients: Web shows collapsible “Thought” rows inside turn folds and filters empty completed reasoning from timeline grouping; mobile drops reasoning in buildThreadFeed. Providers: Claude requests summarized adaptive thinking; Codex turn starts request summary: "detailed" so reasoning items carry text.

Large regression test coverage for ingestion, projection revert, checkpoints, and timeline logic.

Reviewed by Cursor Bugbot for commit 612deb2. Bugbot is set up for automated code reviews on this repo. Configure here.

kgarg2468 and others added 2 commits August 28, 2026 12:28
Provider adapters already normalize thinking deltas to reasoning_text /
reasoning_summary_text, but ingestion dropped them. Reasoning now rides
the assistant message pipeline as an optional channel field: one segment
slot per thread, buffered/streamed with the same delivery switch as
assistant text, stamped with per-thread monotonic timestamps so live and
reloaded order agree, and excluded from answer semantics (turn binding,
checkpoints, titles, search, minimap). Web renders collapsible Thinking
rows; mobile filters reasoning at derivation. Schema change ships as
migration 044, idempotent for databases that predate it.

Design debated with GPT-5.6 sol; implementation by GPT-5.6 sol via Codex
CLI, Claude Opus 5, and Claude Fable 5.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…the timeline

Claude Code 2.1.248+ runs SDK sessions in the redacted-thinking phase:
thinking deltas stream with empty text and only estimated_tokens, so the
reasoning pipeline had nothing to ingest. Pass thinking: { type:
"adaptive", display: "summarized" } (the SDK's --thinking-display
flag) to request API-side thinking summaries, skipped when the thread's
thinking toggle is off.

Debugged and fixed by Claude Fable 5 in Claude Code.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Aug 29, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Team

Run ID: 3bd52164-d610-48c9-a530-2d7961f28cc1

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 29, 2026
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts Outdated
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts Outdated
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts
Comment thread apps/web/src/components/chat/MessagesTimeline.logic.ts Outdated

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web timeline changes for reasoning rows (MessagesTimeline.tsx, MessagesTimeline.logic.ts). The disclosure control matches the file's existing raw-button row pattern (TurnFoldTimelineRow), reasoning rows are correctly kept out of terminal-assistant meta, duration boundaries, minimap previews, and getItemType recycling buckets. One finding on the expanded thought body.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/MessagesTimeline.tsx Outdated
@github-actions github-actions Bot added size:XL 500-999 changed lines (additions + deletions). and removed size:L 100-499 changed lines (additions + deletions). labels Aug 29, 2026
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts Outdated
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding in the web timeline fold logic; details inline.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/MessagesTimeline.logic.ts Outdated
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the timeline fold derivation: reasoning entries now win the "first assistant entry stays visible" slot. Details inline.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/MessagesTimeline.logic.ts Outdated
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts Outdated

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the active-turn "Thinking" indicator. The two items flagged in earlier runs (turn-fold anchoring around a leading thought, and the missing break rule on the expanded thought body) are addressed in the current head.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/MessagesTimeline.logic.ts Outdated

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Effect service conventions review: service definitions, layer wiring, dependency acquisition, and error modeling in this change look consistent with the repository conventions. One minor note on a contracts type that is now duplicated locally.

Posted via Macroscope — Effect Service Conventions

Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts Outdated

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web timeline changes. One finding on the new reasoning disclosure row; the earlier turn-fold, showThinking, and wrap-break-word findings look addressed in this head.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/MessagesTimeline.tsx Outdated
@kgarg2468
kgarg2468 marked this pull request as ready for review August 30, 2026 00:44
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts
Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts
@macroscopeapp

macroscopeapp Bot commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This is a substantial cross-cutting feature that adds persisted provider reasoning, new timeline UI, lifecycle handling, and recovery logic across server and clients. It also changes Claude and Codex provider behavior by enabling detailed reasoning summaries by default, making the product-default impact material.

You can add or adjust custom eligibility rules. Learn more.

Comment thread apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding in the web timeline: the new empty-completed-reasoning skip is applied in the row loop and in deriveTurnFolds, but not before work-entry grouping, so a dropped row can still split an adjacent tool group. Details inline.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/MessagesTimeline.logic.ts Outdated
Comment thread apps/web/src/components/chat/MessagesTimeline.logic.ts

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the new empty-reasoning filtering in MessagesTimeline.logic.ts.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/MessagesTimeline.logic.ts
Comment thread apps/server/src/orchestration/Layers/ProjectionPipeline.ts
Comment thread apps/web/src/components/chat/MessagesTimeline.logic.ts

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 7abafca. Configure here.

Comment thread apps/server/src/orchestration/projector.ts
kgarg2468 and others added 2 commits August 30, 2026 21:42
Codex otherwise emits empty reasoning items when no summary mode is configured. Request detailed summaries on turn/start so the reasoning transcript receives summary text deltas.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@juliusmarminge

Copy link
Copy Markdown
Member

Superseded by #11784 (feat(chat): show provider thinking traces), which landed on main.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL 500-999 changed lines (additions + deletions). vouch:unvouched PR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants