Skip to content

feat(server): preserve budgeted history across provider handoffs - #12352

Merged
juliusmarminge merged 13 commits into
t3code/codex-turn-mappingfrom
t3code/budgeted-context-handoffs
Sep 18, 2026
Merged

juliusmarminge merged 13 commits into
t3code/codex-turn-mappingfrom
t3code/budgeted-context-handoffs

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 18, 2026 •

Copy link
Copy Markdown
Member

Provider switches currently clip each historical message to 240 characters, losing constraints and partial work without bounding the total handoff. This layer above #2829 selects intact historical messages and activity within one budget, preserving roles, order, provenance, and failed/interrupted outcomes. The current request stays separate and unmodified.

Codex injects history before starting one real turn; other adapters, and Codex versions that explicitly reject injection, receive the same selection as attributed context. Model and option changes durably invalidate stale context telemetry before turn start, including failed starts and retries. Codex reasoning-effort-only changes preserve measured usage and capacity for the same model, while discarding the old compaction threshold. Native replacement clears usage while retaining compatible capacity for its handoff budget. Durable delivery records distinguish delivered text from omitted item IDs covered by recovery instructions, preventing repeated handoffs after turn-start failure. An uncertain acknowledgement causes a fresh native thread and context recovery. Omitted history is retrievable through t3_thread_read, now with per-item text offsets beyond its 50,000-character page limit. Existing preview data needs no database migration; V1 transcripts and their recovery excerpts remain available.

The configurable initial cap defaults to 16,000, using conservative UTF-8 byte accounting with a separate 64,000-byte cap on imported history. Native usage, current input, instructions/tools, and subsequent work reduce the context allowance. Budgeting reads live usage from the latest accepted root provider turn, checks its native identity, and uses the model selection that produced that report. Image inputs reserve 8,192 tokens per image instead of charging compressed file bytes against history. Without reported usage, resumed threads also reserve space for previously accepted native attachments. Attempts retain their native thread identity so imported history and replaced threads do not inherit old image charges; legacy recovery coverage supplies the fallback for existing attempts. This fallback uses typical resized image costs from the OpenAI and Claude documentation; dimensions and detail level are not available here, so original-resolution/custom models can differ. Selected-model capacity comes from adapter/catalog metadata or trustworthy telemetry. Unknown windows use a documented 128,000-token fallback; this is not an exact tokenizer or a guarantee for custom models. No additional model call is required.

Validation:

  • 341 focused tests pass across handoff selection/delivery, direct and queued switches, failed/interrupted turns, return/delta, retry recovery, Codex protocol delivery/fallback, merge-back, V1 import, retrieval, wire projection, and contracts.
  • Revalidated on the latest V2 base, including current-input accounting, native/fallback compaction, post-acceptance persistence failure, oversized missed requests, merge-back recovery, and wire redaction. Regressions cover native/fallback switches with one, two, and eight 100 KB screenshots, accounting through the 10 MiB image limit, fresh/model/options capacity changes, failed-start retry telemetry invalidation, and a 7,000-character failed request omitted by a 10,000-character retry followed by a successful 6,000-character ordinary turn.
  • Eight regressions cover prior native images, imported images, reported-usage bypass, unsent inputs, native replacement, and legacy replacement coverage. Existing V1 recovery fixtures verify native attempt identity survives projection rebuild. The prior-image regression fails when the new attachment charge is disabled.
  • Six regressions cover the live image-heavy reasoning-option change, failed-start retry after session reload, compatible capacity on native replacement, small-window constraints, and model/unknown-option invalidation. The live-shaped regression fails before provider start when compatibility handling is disabled.
  • Provider-turn telemetry regressions use the actual running-update/terminal-update event sequence, covering unchanged selections, incompatible model/options with failed retries, and later resumes after native replacement. The Codex protocol fixture checks ownership of real thread/tokenUsage/updated reports. Disabling provider-turn usage lookup reproduces the live failure.
  • Server and contracts typechecks pass. Scoped lint passes with one pre-existing unused-symbol warning in the Claude adapter.
  • Installed Codex CLI 0.154.0's generated protocol includes thread/inject_items; exact request replay tests cover support, method-not-found fallback, and invalid-payload failure.

Independent live verification used actual Codex gpt-6-astra and Claude claude-fable-5-1 on an isolated server. All 12 turns passed: native/inline handoff, two and eight images, return recall, omitted long-request retrieval with paginated text offsets, reasoning-effort change with eight prior plus two current images, and recall after a full server restart. Both providers’ real turn telemetry participates in budgeting. Final follow-ups bound a persistence warning and preserve structured provider start errors. The existing persistence-failure regression and 22 failure/retry tests passed, with server typecheck and scoped lint.

Backend changes only; no UI verification is applicable.

This PR remains directly based on the V2 branch. Native GitHub stack grouping is intentionally omitted because V2 already has a separate child stack; the maintainer chose to keep these parallel changes independent.

Implemented with GPT-6 in the Codex harness.

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XL 500-999 changed lines (additions + deletions). labels Sep 18, 2026
Comment thread apps/server/src/orchestration-v2/ProviderTurnStartService.ts Outdated
Comment thread apps/server/src/orchestration-v2/ProviderTurnStartService.ts Outdated
Comment thread apps/server/src/orchestration-v2/WireProjection.ts Outdated
Comment thread apps/server/src/orchestration-v2/ProviderTurnStartService.ts Outdated
@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.7 KiB — 29.3 KiB ✅
Claude Live turn messages — 1 — 8 ✅

Baseline: unavailable · PR result: bb6da5c · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@macroscopeapp

macroscopeapp Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This change introduces a default-enabled handoff system that selects and delivers historical context across provider switches, retries, compaction, forks, and native protocol sessions. It spans core orchestration, persisted contracts, and provider adapters, with meaningful changes to existing request paths and context budgeting.

You can add or adjust custom eligibility rules. Learn more.

Comment thread apps/server/src/orchestration-v2/ProviderTurnStartService.ts Outdated
@juliusmarminge
juliusmarminge force-pushed the t3code/budgeted-context-handoffs branch from 66ff1dc to 4a12c97 Compare September 18, 2026 02:46
Comment thread apps/server/src/orchestration-v2/Adapters/CodexAdapterV2.ts
Comment thread apps/server/src/orchestration-v2/Adapters/CodexAdapterV2.ts Outdated
Comment thread apps/server/src/orchestration-v2/ProviderTurnStartService.ts Outdated
Comment thread apps/server/src/orchestration-v2/ContextHandoffBudget.ts Outdated
Comment thread apps/server/src/orchestration-v2/ContextHandoffDelivery.ts
Comment thread apps/server/src/mcp/OrchestratorMcpService.ts
@juliusmarminge
juliusmarminge force-pushed the t3code/codex-turn-mapping branch 6 times, most recently from bb072a5 to 1affc9a Compare September 18, 2026 20:19
@juliusmarminge
juliusmarminge force-pushed the t3code/budgeted-context-handoffs branch from 6b22e5d to c790130 Compare September 18, 2026 20:28
@juliusmarminge
juliusmarminge force-pushed the t3code/codex-turn-mapping branch from bbe3dd9 to 01c2162 Compare September 18, 2026 21:28
@github-actions github-actions Bot added size:XXL 1,000+ changed lines (additions + deletions). and removed size:XL 500-999 changed lines (additions + deletions). labels Sep 18, 2026
@macroscopeapp

macroscopeapp Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Macroscope has since reviewed this pull request. An earlier review was skipped by a cost limit; a review has now completed, so that notice no longer applies.

@juliusmarminge
juliusmarminge force-pushed the t3code/budgeted-context-handoffs branch from 6141c79 to 159fc02 Compare September 18, 2026 21:41
@github-actions github-actions Bot removed the size:XXL 1,000+ changed lines (additions + deletions). label Sep 18, 2026
@github-actions github-actions Bot added size:XXL 1,000+ changed lines (additions + deletions). and removed size:XL 500-999 changed lines (additions + deletions). labels Sep 18, 2026
Comment thread apps/server/src/orchestration-v2/ProviderTurnStartService.ts
Comment thread apps/server/src/orchestration-v2/ProviderTurnStartService.ts Outdated
Comment thread apps/server/src/orchestration-v2/ProviderTurnStartService.ts Outdated
@macroscopeapp

This comment has been minimized.

@juliusmarminge
juliusmarminge merged this pull request into t3code/codex-turn-mapping Sep 18, 2026
24 checks passed
@juliusmarminge
juliusmarminge deleted the t3code/budgeted-context-handoffs branch September 18, 2026 23:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL 1,000+ changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant