Skip to content

Providers: tie process lifetime to the prompt-cache TTL #531

Description

@Tryanks

Summary

Tie provider process lifetime to the provider's prompt-cache TTL, and separate it from in-memory log residency. Today a thread's process is kept for 1 h after the user stops viewing it, which is unrelated to what saves money.

How it works today

  • ResidentSessions holds live (viewed) and parked threads (crates/runtime/src/app/sessions.rs). The provider child and the in-memory SessionLog share one lifecycle; leaving residency sends SessionCommand::Shutdown (drop_background, crates/runtime/src/app/lifecycle.rs).
  • Parked idle threads are reaped after RESIDENT_IDLE_GRACE (1 h) with at most MAX_IDLE_RESIDENTS (16) (crates/runtime/src/app/mod.rs). The timer starts at park time, so a thread viewed for 50 min after its last turn keeps its process ~1h50m after that turn. Viewed threads are never reaped.

Facts

  • Anthropic and OpenAI prompt caches are server-side, prefix-matched and keyed per workspace/session, not per process. Claude Code TTL: 1 h for subscription main threads within plan usage, 5 min for API keys, Bedrock/Vertex/Foundry, subagents, compaction (overridable). Codex: 30 min (prompt_cache_options.ttl), prompt_cache_key = thread id, restored on resume.
  • Measured on the maintainer's data (273 Claude threads, 1,425 Codex threads; "miss" = under 50% of previous context read from cache):
Gap between requests Claude, same process Claude, restarted Codex, same process Codex, restarted
< 5 min 0.0% 57% 0.5% 18%
5–60 min 0.2% 71% 8% 27%
1–3 h 91% 100% 38% 75%
> 3 h 86% 100% 100% 96%

Proposal

  1. Split lifecycles. Log residency is bounded by count/bytes and released soon after a thread stops being viewed (cheap to reload with the turn index). Process liveness is governed only by cache and latency value.
  2. Deadline = last turn end + provider TTL, not park or view time. Claude: 1 h or 5 min, detected from the stream's cache-creation breakdown (ephemeral_1h vs ephemeral_5m); Codex: 30 min; others: the provider's documented default. Reap after the deadline, viewed or not.
  3. Count cap ~6–8 live idle processes; evict expired-cache threads first, then least recently active.
  4. Add cache-creation tokens to TokenUsage (crates/agent/src/lib.rs, currently cache reads only). Per Principle 4 it enters the union; providers that do not report it simply lack it. This enables TTL detection and cost display.
  5. Show native cache state in the UI (warm until …), derived from the provider's usage and TTL. Displaying native state fits Principle 1.
  6. No keep-alive pings. They are requests the native CLI would not send (Principle 1).
  7. Flag cache-busting actions: enabling orchestrate, toggling computer-use or changing the model changes the tool list or prefix and forces a full miss on the next request; say so in the UI instead of hiding it.
  8. Fix the comment at RESIDENT_IDLE_GRACE, which states a one-hour cache expiry that only holds for subscription users.
  9. Principle 9: no Claude-version special cases; report findings upstream (#93490, #96101).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestperformanceRuntime performance, latency, resource footprint

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions