Skip to content

Keep the reading position stable while paging long conversations - #454

Merged
Tryanks merged 3 commits into
mainfrom
long-thread-scroll-stability
Sep 16, 2026
Merged

Tryanks merged 3 commits into
mainfrom
long-thread-scroll-stability

Conversation

@Tryanks

@Tryanks Tryanks commented Sep 16, 2026

Copy link
Copy Markdown
Owner

Fixes three reports on a very long (203k-record, 19-turn, pi/deepseek) conversation: visible jumps while scrolling, loading that appears to hang, and the view snapping back to the bottom during continuous up-scroll.

Root causes

  1. Snap to bottom. list_sync fell back to ListSync::Reset (which clears the scroll anchor and re-engages tail following) on 21 of ~1,000 history pages. Two triggers: Timeline::apply_delta looked up a streamed item id across the whole timeline, so when a page revealed an earlier turn, every later turn's stream (pi reuses pi-assistant-0:<i> for each turn) was folded into that earlier turn; and synthetic ids (error-N, compacted-N, …) were numbered from the start of the fold, so any page renumbered later ones. Either way entries changed identity → Reset.
  2. Apparent hang. Pages were 200 records; for a streamed thread a record is a ~5-byte delta, so each page added ~0.1 screen and the thread needed ~1,000 round trips, each re-reading the 25 MB log on the host.
  3. Jumps. The partial first turn changed in place on every page with the reader anchored inside it; and the Prepend re-anchor produced a negative offset_in_item when the reader had scrolled into the history reservation, which gpui leaves unpainted (blank viewport, then a >1-viewport jump on the next wheel notch).

Changes

  • core: apply_delta continues only an entry of the current turn; synthetic_id names timestamped records by timestamp (error-<ts>, -1/-2 on same-ms collisions), untimestamped legacy logs keep the counter. A suffix of the log now folds its turns exactly like the full log.
  • runtime host: the initial snapshot and every page extend back to the record that opens the turn they fall in, inside the unchanged 8 MiB envelope. This thread now loads in 16 turn-sized pages instead of ~1,000.
  • ui: TurnListItem.tail_identity lets Prepend be recognised when the newest loaded row continues; TimelineContinuity { Complete, PartialFirstTurn, NewSession } replaces the reset: bool so a partial first turn is remeasured instead of resetting the list. The Prepend anchor clamps to ≥ 0 and, when the reader was inside the reservation, one on_next_frame walks back by the same distance once the page is measured.
  • docs/DESIGN.md "Timeline history" updated. No new user-visible strings.

Evidence

  • Deterministic replay of the real log through TurnIndexCache::sync as the client pages: before 21 Reset / 14 Prepend / 978 Incremental; after 0 Reset.
  • Real app (debug, macOS) with the thread seeded into a throwaway profile: ~100 up-scroll actions across all 19 turns, 16 pages, all Prepend, position preserved on every page including inside the reservation, no blank frame at screenshot cadence.
  • Tests (each fails with its fix reverted): core streamed_deltas_continue_the_open_turn_not_an_earlier_turn_with_the_same_id, synthetic_entry_ids_do_not_depend_on_how_much_earlier_history_is_folded; runtime history_snapshot_and_pages_start_at_turn_boundaries; ui history_completing_the_only_partial_turn_prepends_instead_of_resetting, completing_the_only_partial_turn_keeps_the_reading_anchor, extended incoming_page_replaces_scrollable_reservation_without_moving_content (asserts a non-negative anchor after the walk-back frame).
  • cargo fmt --all --check, cargo clippy --workspace --all-targets --locked -- -D warnings, cargo build --workspace --locked, cargo test --workspace --locked, cargo machete all green locally.

Known gaps / follow-ups

  • The host still re-reads and parses the whole JSONL per page for a session that never appended in-process (~0.5 s debug / ~0.1 s release per page); caching records changes the never-evicted cache's memory policy, so it is left for a separate change.
  • pi streams under a placeholder id and completes under a different responseId, so a turn can show both the streamed and completed text; that is a crates/agent/src/pi.rs issue outside this fix.
  • The reservation walk-back shows the clamped position for one frame; a synchronous fix needs gpui support for anchors above their row.
  • The reporter's file is a raw per-session event log, not a tcode_thread export, so it cannot be imported through the UI; it was seeded directly into a throwaway profile for testing.

… timestamp

The chat pages history in from the tail, so a suffix of the event log must
fold its turns exactly like the whole log does. Two things broke that:

- apply_delta looked up the streamed item across the whole timeline. pi
  streams every turn's assistant message under the same placeholder id
  (pi-assistant-0:<index>) until completion names it, so a later turn's
  stream was appended to the earliest turn's entry. Live, the text landed in
  the wrong turn; during paging, each page that revealed an earlier turn
  pulled entries out of loaded turns, which the list treated as a
  replacement and reset to the tail. Deltas now continue only an entry in
  the current turn.
- Synthetic ids (error-N, compacted-N, ...) counted from the start of the
  fold, so every page that contained such an entry renumbered all later
  ones. Timestamped records now name themselves; untimestamped legacy logs
  keep the counter.
Streamed conversations hold hundreds of delta records per turn, so a 400
record snapshot and 200 record pages always began mid-turn: the chat folded a
partial first turn whose entries kept shifting as pages arrived, and a 200k
record thread needed a thousand round trips (each re-reading the log) to
load, which read as a hang. The snapshot and every page now extend back to
the record that opens the turn they fall in, within the existing 8 MiB
envelope; the 25 MB / 19 turn repro loads in 16 pages.
Three ways a page could move the reader:

- When the snapshot held a single partial turn, the first page changed that
  turn's identity, so list_sync saw no stable row and either reset the list
  to the tail or appended the new turn at the wrong end. Rows now carry the
  identity of their newest entry, which a page cannot change, and a page is
  a prepend whenever the newest loaded turn continues.
- While earlier history is unloaded, a page can merge the partial first
  turn's streamed fragments so its row shrinks; that was read as a
  replacement and reset the list. The chat passes the history state as
  TimelineContinuity, and the first turn is remeasured instead.
- A reader who had scrolled into the incoming-history reservation was
  re-anchored at a negative offset into the former first turn, which leaves
  the rows above it unpainted until the next scroll and then jumps. The
  anchor now clamps to that row's top and the next frame walks back over
  the measured page, keeping the same pixel position.
@Tryanks
Tryanks merged commit 303a739 into main Sep 16, 2026
6 checks passed
@Tryanks
Tryanks deleted the long-thread-scroll-stability branch September 16, 2026 20:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant