Skip to content

Pi resume of a long thread fails with "Insufficient context allowance for the provider handoff" — same provider, no model change #12929

Description

@astarktc

What happened

I resumed a Pi thread after a couple of days idle (same provider, same model, no settings change — anthropic/claude-fable-5-1, 1M window, ~314k used over 12 turns) and the new turn failed instantly with:

Provider error: Insufficient context allowance for the provider handoff. Compact the target conversation or use a larger-context model; the current request has not been truncated.

I did not switch providers and the model has plenty of headroom, so the advice in the message cannot apply. Every retry fails the same way; the thread is un-resumable in-app.

Diagnosis

Four links, each verified from statev2.sqlite, server.trace.ndjson and a live pi --mode rpc probe on the affected machine:

  1. switch_session exceeds the flat 15 s RPC timeout. Pi re-runs the extension lifecycle on a session switch (MCP servers reconnect, LSP/status extensions re-init). On this project it takes 16.8 s to answer for a 1 MB / 385-entry session file. PI_REQUEST_TIMEOUT_MS = 15_000 in apps/server/src/orchestration-v2/Adapters/PiAdapterV2.ts fires first, so resumeThread fails. (Trace: orchestrationV2.providerTurnStart.start = 15,381 ms with nothing inside but ServerSecretStore.get polls at 5/10/15 s.)

  2. The resume failure is swallowed. ProviderTurnStartService wraps session.resumeThread in Effect.result and takes the provider_resume_fallback branch without logging the cause. The trace has no record of why the resume fell back; I had to reproduce it with a raw RPC probe.

  3. The "fresh native thread" is the old thread. The fallback calls session.ensureThread({ existingProviderThread: { ...providerThread, nativeThreadRef: null } }) — the comment says "the native ref is dropped so the adapter binds a fresh native session instead of retrying the resume that just failed". But PiAdapterV2.registerThread never sends new_session: with a null ref it just runs get_state and adopts whatever session the Pi process is on. Pi does not cancel a switch_session when the caller times out, so by then it had finished switching and get_state returned the original session file. Evidence: the failed run's attempt has nativeThreadId = the original session path, and the provider-thread row still points at it. So sameNativeThread === true for the "replacement".

  4. handoffBudget collapses to 0. The thread's earlier attempts predate feat(server): preserve budgeted history across provider handoffs #12352 and carry no nativeThreadId, so latestNativeContextUsage cannot attribute the reported usedTokens 314551 / maxTokens 1000000 → reportedUsage is null → the window falls back to the 128,000 default (PiAdapterV2 implements no getModelContextWindow, although Pi's get_state reports model.contextWindow = 1000000). With sameNativeThread true, nativeContextEstimate sums the saved history (~307 KB here) → 128000 − 307543 − 30 − 32000 < 0 → budget 0 → not even the coverage marker fits → ContextHandoffBudgetError (ContextHandoffDelivery.ts).

Net effect: a spurious timeout triggers a handoff from the thread to itself, budgeted as if the model had 128k and the whole history were foreign. Had the switch answered in 14 s, no handoff would have been prepared at all.

Suggested fixes (independent; A–C would close it):

  • A. PiAdapterV2: honor the fresh-thread contract. When existingProviderThread.nativeThreadRef === null and this process has already attempted a switch_session, send new_session before get_state (Pi RPC supports it). Guard with a flag so the ordinary first-turn path pays no extra lifecycle cost.
  • B. Give switch_session a lifecycle-sized timeout (e.g. 60 s) instead of the generic 15 s.
  • C. Log the resumeThread failure cause before falling back (Effect.tapError → Effect.logWarning), so the trace explains the fallback.
  • D. Implement getModelContextWindow in PiAdapterV2 from the get_state contextWindow, so Pi never lands on the 128k default when telemetry is unattributable.
  • E. Orchestrator guard: if the replacement's nativeThreadRef equals the original's, the adapter did not replace the thread — fail with a contract error (or treat as resumed) rather than budgeting a self-handoff.

Steps to reproduce

  1. Pi provider, a project with a few extensions/MCP servers — enough that Pi's switch_session takes > 15 s. Check from the project cwd:
    printf '{"type":"switch_session","sessionPath":"<session file>","id":"sw"}\n' | pi --mode rpc and time the sw response.
  2. Run a thread for a dozen turns so the saved history is a few hundred KB. (Attempts created before feat(server): preserve budgeted history across provider handoffs #12352 make it deterministic, but any long thread is enough once nativeContextEstimate exceeds ~96k bytes under the 128k default.)
  3. Let the Pi session close (restart the app / idle out), then send another message to the thread.

Expected: the native resume succeeds, or — if it genuinely fails — the fallback binds a new Pi session as the code comment promises and the handoff is budgeted against the model's real window.
Actual: run fails in ~150 ms with ContextHandoffBudgetError; every retry replays the same path.

Version

0.0.42 — fork of the Orchestrator V2 branch (PR #2829) at b488c57, Electron 44.4.2; the resume/handoff code cited is upstream-unmodified

Environment

macOS 26.6 (arm64), desktop app against its local server, Pi 0.85.1, model anthropic/claude-fable-5-1 (thinking high), full-access runtime

Evidence

server.trace.ndjson (times relative to the turn-start span):
  orchestrationV2.providerTurnStart.start           durationMs=15381.7
    PiAdapterV2.openSession                          +0 ms
    ServerSecretStore.get                            +5 s, +10 s, +15 s   (nothing else in the window)
    orchestrationV2.contextHandoff.prepareProviderHandoff   +15.14 s
    orchestrationV2.deliverContextHandoffs           Failure: ContextHandoffBudgetError: Insufficient context allowance for the provider handoff. Compact the target conversation or use a larger-context model; the current request has not been truncated.
  "orchestration V2 provider turn start failed" { _tag: "ProviderAdapterTurnStartError", driver: "pi", cause: { _tag: "ContextHandoffBudgetError" } }

statev2.sqlite after the failure:
  orchestration_v2_projection_context_transfers: type=provider_resume_fallback, sourceThreadId == targetThreadId, status=resolved_portable
  orchestration_v2_projection_context_handoffs:  strategy=full_thread_summary, summaryText "Selected 8 intact items; omitted 206 items"
  orchestration_v2_projection_run_attempts (failed run): nativeThreadId = the ORIGINAL session file (same one the previous 12 runs' provider thread points at)
  orchestration_v2_projection_provider_turns (last completed): tokenUsage usedTokens=314551 maxTokens=1000000; earlier attempts have no nativeThreadId
  user message of the failed run: "$handoff" (no attachments)

Live probe, `pi --mode rpc` from the project cwd (times from stdin write):
  10.22s extension_ui_request setStatus pi-lens-lsp
  14.29s extension_ui_request setStatus mcp "MCP: 2 servers enabled"
  16.82s response id=sw command=switch_session success=true {"cancelled":false}
  (get_state on the same process reports model.contextWindow=1000000)

Related issues

#12352 introduced the budget path and the attempt nativeThreadId that link 4 depends on. No other match for "Insufficient context allowance" / provider_resume_fallback in open or closed issues.

Fix applied or workaround

Nothing was changed on the affected machine. No in-app recovery exists: every retry replays switch → timeout → self-handoff → budget 0. Options: resume the session from a terminal (pi --resume <session file> in the project cwd), or start a new thread and pull history via t3_thread_read. Locally I am carrying a fork patch for A+B (60 s lifecycle timeout for switch_session/new_session; new_session before get_state when a fresh thread is requested after a switch/fork on the same process) — happy to open it as a PR against the V2 branch if wanted.

Filed by

claude (anthropic/claude-fable-5-1) running as a Pi thread inside T3 Code, following the t3 triage playbook

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions