Skip to content

Feature Request: Automatic Context Window Management / Trimming #114

Description

@james-333i

Problem:
Foundation Models in AnyLanguageModel have a hard context window limit (e.g., 4096 tokens). Currently, when this limit is exceeded, the session fails with an error:
exceededContextWindowSize(FoundationModels.LanguageModelSession.GenerationError.Context(debugDescription: "Content contains 4098 tokens, which exceeds the maximum allowed context size of 4096.", underlyingErrors: [Provided 4,098 tokens, but the maximum allowed is 4,096.]))

Other models typically have higher limits but will reach a limit at some point as well.

While this is something the developer could handle, since this library is treated as universal and backends like OpenAI, Claude, etc. often handle this kind of adjustment automatically, it would be ideal for the library to tackle this for local models.

This is not just a problem for long conversations — tool calls can also contribute large amounts of data. For example:
• Tool outputs containing structured data, summaries, or long document content
• Multiple tool calls within the same session, each adding their prompt and result
• Reference guides, lookup tables, or document embeddings stored in session history

All of these count toward the token limit, making it easy to exceed 4096 tokens even if the visible chat is short.

Proposed Enhancement:
1. Token Tracking per Session
• Automatically count tokens for every piece of context in a session: system prompts, user messages, assistant responses, tool calls, and tool outputs.
2. Automatic Trimming / Summarization
• When adding new content would exceed the model’s context window:
• Drop or summarize the oldest messages and tool outputs until the new input fits.
• Optionally allow developers to mark certain tool outputs or messages as “persistent” so they are never discarded.
3. Tool-Aware Handling
• Treat tool calls specially:
• Tool inputs and outputs may be large but can be summarized or compressed when stored in session history.
• For example, storing only essential fields or summaries instead of full JSON.
4. Configurable Strategy
• Developers could define custom trimming rules:
• Drop entire messages
• Summarize tool outputs
• Retain only the last N tool calls
• Compress historical context into a short summary
5. Unified Across Models
• Apply this strategy for all model types (Foundation, MLX, etc.), respecting their individual max token limits.

Benefits:
• Prevents exceededContextWindowSize errors without manual intervention.
• Makes long-running sessions with multiple tool calls robust.
• Enables developers to safely use heavy reference data, guides, or tool outputs without exceeding the model limit.
• Provides a clear framework for context management that can scale as models with larger windows are added.

Activity

  1. james-333i commented on Feb 8, 2026

    @james-333i
    ContributorAuthor

    Adding this other library I found for reference. I did try to see if I could get it to work with AnyLanguageModel and with MANY hacks I did but would be much more practical to build in something like this directly in to AnyLanguageModel. This other library I had to mess with minimum SDK version, replace all FoundationModels imports with AnyLanguageModel and extend missing functions.

    https://github.com/Silo-Labs/swift-context-management

    The concept, however, is good. It supports reductions like:
    HeadTailWindowReducer
    HierachicalSummaryReducer
    RollingSummaryReducer
    SlideringWindowReducer
    etc.

    To do this well, you need to keep track of token count. Some models may provide this automatically but others may need a library like Tiktoken to count. Swift tiktoken could be an option.

  2. christopherkarani commented on May 18, 2026

    @christopherkarani
  3. mattt commented on Sep 14, 2026

    @mattt
    Collaborator

    This depends on contextSize and tokenCount(for:), which #206 brings in from the OS 27 API, so I'm treating it as a 2.0 item under #210. The shape I'd like is a session-level policy rather than per-provider logic. Leaving open.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions