Skip to content
This repository was archived by the owner on Sep 8, 2026. It is now read-only.
This repository was archived by the owner on Sep 8, 2026. It is now read-only.

plan: A5 viewport — phase 1 — bounded best-effort tail read and snapshot-first attach #959

Description

@btipling

Plan header

Field Value
Status HANDOFF-READY
Date 2026-09-07
Type phase 1 — one issue / one PR
Parent #958 — tracking only
Source #553; user-facing bug #924 closes in phase 2, not here
Branch main baseline b31ca512e7ddee162fa8313cc52f549822e2d23a; proposed feat/a5-bounded-tail-read
Layers Vercel backend + browser-safe types/parser; no production host flip or Wasm export changes
Reusability Existing injected session/Blob/Workflow adapters, no deployment ids or new secret
Production administrative mutate? No — read-only recovery, no migration/backfill/env cutover
Cloud path Existing Git deploy, int-durable and current Wasm artifact
Living docs docs/session-model.md, docs/agent-stream.md, docs/harness-limits.md, SECURITY.md, AGENTS.md

Review notes — decision and unblock (2026-09-07)

Operator decision, this session: “we don't need perfect precision, best effort is good enough. default to good perf and sensible defaults … History is a nice to have … Pick reasonable defaults, don't block on open questions, resolve them.”

This body replaces the previous precision-first design. The former L1/L2 blockers are resolved by a simpler product contract: bounded recent display, explicitly incomplete history, and no attempt to reconstruct exact old-run boundaries. This is intentional omission/approximation in a disposable viewport, not permission to bypass authentication or overwrite the canonical transcript with it.

Findings resolved

Severity Previous issue Resolution
Blocker, resolved Legacy prompt/run splice cannot always be proven Do not reconstruct it. Display recent stream rows or latest stored head, with an incomplete-history note. Missing prompt is acceptable.
Blocker, resolved Exact overlapping-chain matching is expensive No chain walk or overlap matcher in cold recovery. One stored head and a bounded stream sample only.
Major, resolved Catch-up may chase a fast producer forever One bounded sample and one final tail probe. Skip unsampled/backlogged frames to that tail with an explicit gap flag.
Major, resolved Certificates/witnesses require new persistence machinery Deleted from scope: no displayResume, hashes, step witnesses, encrypted run-input reads or producer metadata. A resume index is a transport position, never proof of complete history.
Major, resolved Partial viewport could truncate cloud history Response has historyComplete:false; phase 2 must enforce no full-snapshot PUT from this source. This phase makes no application writes.

Scores after edits: correctness 4, performance 5, architecture 4, testing 4, security 4, reusability 5, parent adherence 5, layers 5, cap governance 5, living docs 5. Cloud administrative ops N/A (none); layout/palette N/A (no UI change). Verdict: HANDOFF-READY. No tests or application edits executed while reviewing. Read-only baseline and SDK APIs verified in this session. All questions in this phase are decided below.

Intent and goals

In: bounded latest-head reader, bounded recent-stream reducer, snapshot-first cold GET mode, explicit absolute cursor framing, pure parser/types, tests/docs.

Out: changing worker persistence, inference/A4, exact history fidelity, full-chain recovery, global CAS, queue persistence redesign, SDK upgrade/dependency addition, historical-run browser APIs, new producer segment ids, new host UI, cache/scrollbar implementation.

Goal Observable success
G1 A large live run can be attached without reading from origin or shipping its historical thinking to the browser
G2 Recovery work is bounded independently of total run length; budget exhaustion still allows tail attach
G3 Response distinguishes sampled rows, omitted data, transport resume position and run status
G4 Current unnegotiated clients and inference/persistence behavior are unchanged
G5 Unauthorized/foreign session/run access fails closed and the real-Wasm foundation int rows pass

Architecture and verified baseline

Concern Owner / paths
Scoped read service Backend: new lib/sessions/viewportRead.ts; injected stores from lib/di/index.ts
Bounded Workflow read adapter Backend factory: new lib/workflows/viewportRunReader.ts; public workflow/api.getRun, no private runtime imports
Projected view/reducer New browser-safe lib/sessionViewport.ts + tests; existing lib/toolRun.ts encoding
Snapshot/indexed transport Existing app/api/turns/[runId]/stream/route.ts, app/api/turns/route.ts, new lib/agent/viewportStream.ts
JSON latest head New app/api/sessions/[id]/viewport/route.ts + route tests
Versioned protocol parser New lib/viewportStreamProtocol.ts + tests; no host wiring until #960
Test integration int/viewport-read.int.test.ts (new), existing int/{driver,stores,loadBridge}.ts

Verified on baseline:

  • coldAttachFromSnapshot and cold Send path hardcode zero; harnessChat wipes through last user; turnApply always paints reasoning and bumps a callback count. Host flip belongs to plan: A5 viewport — phase 2 — usable tail-first F5 and real-Wasm #924 close gate #960.
  • Pinned workflow/core 4.8.4 public getRun().getReadable({startIndex}) returns decoded stored frames and has getTailIndex() (-1 empty). Use absolute start after a captured tail, not a negative index resolved at an unknown later moment. SDK handles encryption internally.
  • Existing GET stream auth verifies owned session envelope and meta.turnRunId===runId before reading. Reuse that current-run check. No arbitrary historical run ids accepted by this API.
  • Existing stream wrapper distinguishes stored events from synthetic terminal status; new framing must do the same.
  • Blob heads have messages, optional queue, prev/depth; running heads can be overlapping this-run snapshots, terminal heads usually flatten. Cold view is head-only, never calls reconstructTranscriptChain or mergeCheckpointOntoPrior.
  • Blob scope validation isObjectIdBoundTo already exists. Current BlobTranscriptStore.read returns one whole body; the existing object rail is 8 MiB. This plan does not claim byte-range I/O or exact heap = wire bytes.
  • int-durable runs npm run test:int; int/loadBridge loads real current Wasm and fails closed. No new native source needed.

Locked decisions

1. Authorization and finite optional sources

Authenticate using existing session/tenant gates; require owned session. For live read/attach require requested run matches current envelope run before SDK access. Reject 400 invalid syntax/version, 401 unauthenticated, 404 foreign/mismatched/missing scope, 503 when mandatory session store or initial live-tail metadata is unavailable. No keys/raw run inputs/signed URLs in responses/logs.

Only read the currently authorized transcriptPointer for this phase; bind object id to session. One Blob read, no prev traversal. A missing/corrupt/too-large head is an optional-history miss, not an excuse to replay from zero or cancel a live run. Validate row roles/text; ignore unknown Blob metadata, especially any old proposed certificate fields. Return only safe carriers: selectedModel/reasoningEffort/provider/cwd/bind/persona id/attached slugs/usage/run state and sanitized host-known queue. Never return arbitrary meta, personaSnapshot, notes/model/compaction bodies.

For the JSON head route, return head tail rows under view caps and hasEarlier: boolean from remaining rows or a valid prev, but do not claim global history completeness. When no usable head exists, return a typed empty/unavailable view with replace:false; #960 retains cached paint. No live stream for ordinary idle/terminal JSON restore.

2. Cold sample algorithm — no exact splice

The negotiated GET is one response, not JSON then a later raw GET:

  1. After authorization emit viewport_state with actual run id/liveness, so plan: A5 viewport — phase 2 — usable tail-first F5 and real-Wasm #924 close gate #960 can show Busy/Stop while recovering.
  2. Capture H0 = SDK tailIndex+1. Start recovery clock when optional history work starts. Read at most one head and the bounded stream sample (can run concurrently under the same deadline).
  3. Sample decoded frames from S=max(0,H0-2048) to H0, stopping at the first of: 2048 frames, 8 MiB decoded frame bytes, 5 seconds recovery work, client abort, EOF/error. No origin scan when H0 exceeds 2048. No source refetch/retry. Discard reasoning text immediately. Work and retained reducer state are bounded separately; don't collect an array of all sample events.
  4. Prefer sampled assistant/tool/error display rows if any exist. Otherwise use bounded latest head rows. Choose one source, do not heuristically concatenate head and sample. This avoids expensive/dangerous duplicate alignment. Missing historical user/prompt/earlier rows is acceptable. If both empty, emit replace:false so cached paint remains, with a neutral “Attached to live run; earlier history unavailable” note. Do not label this missing data a user/model error.
  5. Cancel/release the sample reader. Make one fresh tail probe (maximum wait 1 second) and set H=max(H0,freshTail+1). If that probe fails/times out, use H0. Never re-probe in a loop. Frames omitted because of sample limits, errors or growth through H are intentionally skipped and incomplete/gap flags are set. Initial H0 is required; without a trustworthy transport position fail the attach, keep cached view, offer retry.
  6. Emit one viewport_snapshot with bounded rows, source/approximation flags and resumeIndex:H. Open/continue the live decoded stream at exactly H on this same response. Post-H frames are live relative to this server handoff, even if network delivery is delayed. No proof that every pre-H display byte is present is claimed or needed.

The 5-second budget is for optional recovery, not a new timeout on the live Workflow or existing long-lived stream route. Waits are abort-raced; cancel readers on budget/abort, stop issuing work, observe late promise errors and never let late results replace the chosen view. SDK/Blob calls without abort support may finish in the background; only the finite launched calls may do so, no continuing app loop. One decoded oversized SDK frame may already have allocated before budget detection; discard it and skip, no attempt to reserialize/copy it repeatedly. Document this honest bound rather than implementing custom SDK byte decoding.

3. Simple display reducer / incomplete data

  • Rows are plain projected display messages, never a full canonical SessionSnapshot. Clip oversized row display to existing 262144-byte bridge ceiling with a visible excerpt marker; source Blob untouched.
  • Adjacent sampled text deltas concatenate into a bounded recent assistant excerpt. On prefix loss, mark it as a continuation; do not retain a hidden full assistant string.
  • Pair tools by call id when available within the sample. Missing start produces a completed result row; missing result leaves a sampled pending item labeled incomplete. Fall back to separate items when identity is ambiguous; don't attempt exact execution dedup. Reuse existing tool group limits and encoding; don't mutate a source row outside the sample.
  • done.text is a full-run aggregate. For sample display use its bounded tail only if there is no sampled assistant text; otherwise do not append it. Error/terminal rows can be shown without claiming a complete transcript.
  • All sampled reasoning_delta bodies are omitted. All post-H reasoning deltas may paint as new live thinking, including continuation of an old producer segment: the handoff is the sensible boundary. Exact old adjacent-segment recovery is explicitly waived by the operator's best-effort decision. No new producer fields or inference code needed.
  • The host will reset live assistant/tool/thinking buffers after snapshot and begin fresh continuation rows rather than rewrite an unrelated restored head row. Modest duplicate/fragmented display is accepted; no full replay, no runaway buffers.
  • Always expose historyComplete:false, incomplete:true, source (stream_tail/stored_head/unavailable), sampled range and skipped/gap flags. Mark omissions in one concise System note, not one row per dropped token. No exact total history count or prior/current-run splice.

4. Negotiated wire

Existing GET query modes:

  • viewportVersion=1&hydrate=tail — cold sample+snapshot+live; startIndex with this mode rejects.
  • viewportVersion=1&startIndex=N — hot indexed resume, no bootstrap; caller has same-heap state.
  • no version — original behavior untouched.

Durable POST may negotiate ?viewportVersion=1 to use indexed events starting at zero for a genuinely new live run. Unknown/duplicate/conflicting selectors reject. There is no production cold consumer in this phase.

SSE records (browser-safe discriminated parser):

event data Cursor
viewport_state {version:1,runId,status,phase:'recovering'} none
viewport_snapshot {version:1,runId,resumeIndex:H,replace,rows,source,historyComplete:false,incomplete:true,gap,sampledRange,carriers} establish H explicitly after snapshot handling; gaps are intentional and flagged
turn_event {version:1,runId,nextIndex:I+1,event:AgentStreamEvent} + SSE id:I+1 absolute raw index after application
viewport_end {version:1,runId,status} from synthetic terminal status none; not transcript-completeness evidence
viewport_error {version:1,runId,code} none; preserve last applied position

Use one decoded stored frame per raw index; fragmenting network bytes does not increment. Synthetic records have no stored id. Malformed stored frame consumes its known raw index and emits {version:1,runId,nextIndex:I+1,skipped:true} under turn_event with matching SSE id (no event property), never untrusted content; transport/decoder failures that make the raw position unknown close the reader, not guess the cursor. Parser validates monotonic indices and id/data agreement; only explicit snapshot may jump. Existing event payload parsing/redaction remains.

Use existing stream content type and no-buffer headers, add private,no-store,no-transform. Close/cancel reader and polling on disconnect/terminal/error. Reuse existing poll cadence and terminal classification; separate stored done/error from synthetic status. A terminal run can still return sampled/head display; never wait for an exact terminal Blob commit or reread a chain. No application persistence, inference, run start or run cancel during read recovery.

5. Resource fitting

VIEWPORT_RESPONSE_MAX_BYTES=2 MiB measures full escaped snapshot/JSON payload including carriers and framing overhead reserve. At most existing 2048 rows, drop oldest projected rows until fit, newest row excerpt if needed, one omission note. Per-row projection uses existing bridge byte ceiling. Budget object construction incrementally (row-byte accounting), not repeated serialization of full source arrays. Entire live SSE can exceed 2 MiB over time; the limit is an individual snapshot/control response, not the stream lifetime.

No growth proportional to run history: reducer retains a bounded output window, current text excerpt and bounded tool group only. Release head raw JSON/unused messages after selection. Complexity is linear in one legal Blob body plus capped sampled frame bytes, not number of all historical objects or events. No reference ropes, suffix matcher, hashes, witnesses, step-input decoding or certificate maintenance.

Caps and operator authorization

The operator explicitly authorized reasonable performance-first defaults and best-effort loss/omission on 2026-09-07. These are NEW recovery/view budgets; no existing durable storage, inference, step, queue, bridge or live-route cap changes. Existing-cap change still needs a separate decision.

Cap Value Defense / location
NEW VIEWPORT_TAIL_MAX_FRAMES 2048 One generous recent sample, constant work regardless of multi-hour run; lib/sessionCloudCaps.ts
NEW VIEWPORT_RECOVERY_MAX_BYTES 8 MiB decoded sample bytes Generous display recovery comparable to existing object ceiling; includes reasoning bytes even though discarded; caps reducer input work
NEW VIEWPORT_RECOVERY_MAX_MS 5000 ms Optional history cannot hold F5 for minutes. Live inference/attach unaffected. Common bounded sample should complete sooner.
NEW VIEWPORT_FINAL_PROBE_MAX_MS 1000 ms One final metadata refresh then H0 fallback; no chase loop. Derived from existing status-poll scale, not altered SDK retry setting.
NEW VIEWPORT_HEAD_READ_MAX_OBJECTS 1 No automatic ancestor traversal on cold path; not lowering existing 256-object reconstruction cap elsewhere
NEW VIEWPORT_RESPONSE_MAX_BYTES 2 MiB Same safe Function-body ceiling, <actual 4.5 MB; full escaped payload, not 2048×row assumption
Existing output rows / row bytes 2048 /262144 Reuse ring/bridge; explicit display excerpts, no source retention change
Existing object / meta / Function 8 MiB /1 MiB /2 MiB Unchanged. One Blob read may allocate raw+parsed overhead, no claim heap equals bytes
Existing tool groups 200 items /262144 encoded B /229376 detail budget Stream group rollover, not all-run tool state
Existing raw cursor / poll 1000000000 /1000 ms Same transport/liveness rails
Existing turn/route 1-hour turn, 300000-ms wall wrap, 512 steps; stream maxDuration1800 s Unchanged, no per-token steps
Source concurrency One head read + one sample reader, at most initial and final tail probes Finite launched work; cancellation/late-result guards

Testing — minimum shipping matrix

No tests run during plan review. All delivered rows ordinary green it, not skipped/expected failure.

# Case Layer / assertion
T1 Large lazily generated run, H0≫2048, historical reasoning dominates Adapter/service: first sample starts at max(0,H0−2048), ≤2048 frames, no origin/step/run-input/prev reads; no reasoning in snapshot
T2 Sample ends on frame/byte/time budget; head missing/corrupt; final probe slow/fails; constant producer growth Controlled clocks/IO: finite work, exactly one final probe, gap flags, H0 fallback, live attach rather than catch-up loop
T3 Old run with no prompt, repeated prompt, ambiguous tools/segments, current turn before first persist Successful best-effort tail or replace:false, no certification requirement; future post-H reasoning works
T4 Select stream display vs head fallback; 8 MiB head with valid prev One head only, no mixed-source exact merge; source immutability by zero writes
T5 Unicode/escaped JSON/oversized last row/full queue, tool partials/done aggregate Actual payload≤2 MiB, projected row≤bridge cap, no whole-run strings/duplicate done append; explicit omissions
T6 Unauthorized/foreign session/run/Blob, missing store, request deleted/switched before handoff Fail closed before forbidden source access; no secrets or arbitrary metadata returned; no start/cancel/write
T7 Version negotiation, old GET/POST parity, raw-index/id framing, split UTF-8/CRLF, synthetic terminal/error, unknown position Parser/route contract; no silent guessed cursor; default old clients unchanged
T8 Empty/running/cancelling/completed/failed/cancelled run, disconnect/budget/EOF reader cleanup Liveness independent of complete history; no orphan app loops or automatic inference cancel
T9 Read service→real protocol parser→real Wasm ring New int/viewport-read.int.test.ts; newest sampled rows, absent historical thinking, omission note, subsequent live text/thinking; mock external IO only
T10 All gates npm test, npm run test:int, npm run typecheck, npm run build; current Wasm artifact and DI/staticGraph gates

#924 requires #960's production-controller real-Wasm tests; T9 is foundation evidence, not a premature close claim. Existing cold-zero helper tests remain until host phase changes them. Resource tests assert deterministic call/byte/frame budgets and use lazy fixtures, not a flaky wall-clock microbenchmark.

Living docs and cloud ops

Surface Change
docs/session-model.md Recent best-effort read, incomplete history, no destructive restore write; default host not yet flipped
docs/agent-stream.md Negotiated grammar, transport cursor not full-history proof, snapshot/live boundary and reasoning policy
docs/harness-limits.md New recovery budgets/omission behavior and honest one-frame/SDK-allocation residual
AGENTS.md Where to change bounded viewport; don't resurrect exact certificate/merge prerequisites; real-Wasm gate
SECURITY.md Current-run/session scope, read-only recovery and no credential/raw-input exposure
README N/A in this phase: no user-facing host flip; phase 2 owns feature claim
.env.example, native protocol guide N/A: no env/secret/exports/version change

Use existing Git deploy and int-durable current artifact. No Production migration, setup nag, new service, SDK enable toggle or laptop step. Rollback code only; storage/producers untouched. Docs describe current behavior, not phase/issue archaeology.

Implementation order / DoD

  1. Add caps, view types/pure reducer, protocol parser and tests.
  2. Add injected bounded head/Workflow adapters and read service; security/budget tests.
  3. Add negotiated GET cold/indexed and POST indexed modes, keep default behavior; terminal/cleanup tests.
  4. Add JSON latest-head route, real-Wasm foundation int, docs; run required gates. One PR, no merge under implementation skill.

Parent refinement / resolved questions

This applies the operator's new best-effort priority. Former exact legacy splice, cryptographic checkpoint, source-span matcher, unknown segment and catch-up-proof gates are removed, not moved to another mandatory phase. Later history paging is bounded manual/boundary exploration, not perfect reconstruction. No user decisions or design blockers remain here.

Residual risk: recent rows may be fragmented, duplicated, missing or briefly stale; a burst produced during sampling may be skipped; source budgets may fall back to stored/cached paint. Those are accepted display tradeoffs. Authentication, session isolation, no accidental transcript replacement and green real-Wasm budget/no-replay tests remain merge gates.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions