feat(server): preserve budgeted history across provider handoffs - #12352
Conversation
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
ApprovabilityVerdict: Not approved Macroscope's review found this PR not approvable — This change introduces a default-enabled handoff system that selects and delivers historical context across provider switches, retries, compaction, forks, and native protocol sessions. It spans core orchestration, persisted contracts, and provider adapters, with meaningful changes to existing request paths and context budgeting. You can add or adjust custom eligibility rules. Learn more. |
66ff1dc to
4a12c97
Compare
bb072a5 to
1affc9a
Compare
6b22e5d to
c790130
Compare
bbe3dd9 to
01c2162
Compare
|
Macroscope has since reviewed this pull request. An earlier review was skipped by a cost limit; a review has now completed, so that notice no longer applies. |
6141c79 to
159fc02
Compare
Provider switches currently clip each historical message to 240 characters, losing constraints and partial work without bounding the total handoff. This layer above #2829 selects intact historical messages and activity within one budget, preserving roles, order, provenance, and failed/interrupted outcomes. The current request stays separate and unmodified.
Codex injects history before starting one real turn; other adapters, and Codex versions that explicitly reject injection, receive the same selection as attributed context. Model and option changes durably invalidate stale context telemetry before turn start, including failed starts and retries. Codex reasoning-effort-only changes preserve measured usage and capacity for the same model, while discarding the old compaction threshold. Native replacement clears usage while retaining compatible capacity for its handoff budget. Durable delivery records distinguish delivered text from omitted item IDs covered by recovery instructions, preventing repeated handoffs after turn-start failure. An uncertain acknowledgement causes a fresh native thread and context recovery. Omitted history is retrievable through
t3_thread_read, now with per-item text offsets beyond its 50,000-character page limit. Existing preview data needs no database migration; V1 transcripts and their recovery excerpts remain available.The configurable initial cap defaults to 16,000, using conservative UTF-8 byte accounting with a separate 64,000-byte cap on imported history. Native usage, current input, instructions/tools, and subsequent work reduce the context allowance. Budgeting reads live usage from the latest accepted root provider turn, checks its native identity, and uses the model selection that produced that report. Image inputs reserve 8,192 tokens per image instead of charging compressed file bytes against history. Without reported usage, resumed threads also reserve space for previously accepted native attachments. Attempts retain their native thread identity so imported history and replaced threads do not inherit old image charges; legacy recovery coverage supplies the fallback for existing attempts. This fallback uses typical resized image costs from the OpenAI and Claude documentation; dimensions and detail level are not available here, so original-resolution/custom models can differ. Selected-model capacity comes from adapter/catalog metadata or trustworthy telemetry. Unknown windows use a documented 128,000-token fallback; this is not an exact tokenizer or a guarantee for custom models. No additional model call is required.
Validation:
thread/tokenUsage/updatedreports. Disabling provider-turn usage lookup reproduces the live failure.thread/inject_items; exact request replay tests cover support, method-not-found fallback, and invalid-payload failure.Independent live verification used actual Codex gpt-6-astra and Claude claude-fable-5-1 on an isolated server. All 12 turns passed: native/inline handoff, two and eight images, return recall, omitted long-request retrieval with paginated text offsets, reasoning-effort change with eight prior plus two current images, and recall after a full server restart. Both providers’ real turn telemetry participates in budgeting. Final follow-ups bound a persistence warning and preserve structured provider start errors. The existing persistence-failure regression and 22 failure/retry tests passed, with server typecheck and scoped lint.
Backend changes only; no UI verification is applicable.
This PR remains directly based on the V2 branch. Native GitHub stack grouping is intentionally omitted because V2 already has a separate child stack; the maintainer chose to keep these parallel changes independent.
Implemented with GPT-6 in the Codex harness.