Skip to content

fix(server): the Claude 200k context window selection takes effect - #14988

Open
Mnigos wants to merge 2 commits into
pingdotgg:mainfrom
Mnigos:claude-200k-context-window
Open

Mnigos wants to merge 2 commits into
pingdotgg:mainfrom
Mnigos:claude-200k-context-window

Conversation

@Mnigos

@Mnigos Mnigos commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #8405.

Problem

Claude Code runs a 1M-capable model at 1M when given its bare slug, unless CLAUDE_CODE_DISABLE_1M_CONTEXT is set. T3 encoded the context window only in the slug ([1m] for 1M, the bare slug for 200k), so selecting 200k did nothing and the session still ran at 1,000,000 tokens. On main the composer meter and the stored tokenUsage.maxTokens come from the selected option, so they read 200k while Claude Code reported a 1,000,000-token window for the same turn. Separately, claudeContextWindow hardcoded 1M for claude-opus-4-6, which has a 200k/1m selector, so its meter showed 1M with 200k selected.

Change

  • ClaudeModelCatalog.ts: new resolveClaudeCatalogContextWindowEnv. For a model whose catalog entry has a context window selector, it returns CLAUDE_CODE_DISABLE_1M_CONTEXT: "1" when the selected (or default) window is 200k or less and "0" otherwise. The window is stated both ways so the selection, not an env in the user's Claude settings, decides it; Claude Code also refuses a [1m] slug while the variable is set. Models without a selector (Opus 4.7, Haiku 4.5, custom models) get no value and keep whatever the user configured.
  • claudeModelOptions.ts: compileClaudeModelSelection returns that env and includes it in queryIdentity. The adapter reuses a live query only when the identity matches, and Claude Code reads the env when it starts, so switching between 200k and 1m on a thread opens a new query, also for a catalog model whose API id does not change with the window.
  • ClaudeAdapterV2.ts: makeClaudeQueryOptions passes the env as settings.env on the SDK query, next to the selection's other settings. These are flag settings, which outrank env from the user's settings files. The existing merge with the instance's SDK settings and autoCompactWindow is unchanged and keeps the env.
  • ClaudeAdapterV2.ts: claudeContextWindow no longer hardcodes 1M for claude-opus-4-6. Opus 4.6 has a 200k/1m selector in the catalog, so its stored maxTokens and meter now follow the selection like the other selector models. claude-opus-4-7 has no selector and stays at a fixed 1M.

Scope and approval

Triaged bug #8405, following the triage comment, which adds the acceptance case from #5286 (start a new Claude session, select 200k, read the meter): "After the provider fix, this sequence must make both the live provider window and the meter show 200k." Server and Claude adapter only; no contract or client change, since the clients already render the stored maxTokens. The recorded Claude replay transcripts (41 files under orchestration-v2/testkit/fixtures) now expect settings.env.CLAUDE_CODE_DISABLE_1M_CONTEXT: "1" on their 59 query.open frames, because they were recorded on claude-sonnet-4-6, whose catalog default is 200k.

One behavior change for selector models: with no explicit selection, the catalog default is now stated (Opus 4.6 defaults to 1M and gets "0", Sonnet 4.6 defaults to 200k and gets "1"), so a CLAUDE_CODE_DISABLE_1M_CONTEXT in a user's Claude settings no longer decides the window for these models; the composer selection does. Models whose default is 200k (Sonnet 5.5, Sonnet 5, Sonnet 4.6) used to run at 1M on a new thread while the meter showed 200k; they now run at the 200k the meter shows, and 1M is an explicit selection.

Not changed: ClaudeTextGeneration (titles, commit messages) still passes only the model id to its one-shot claude -p calls.

Verification

Observed result. Real web client, headless Playwright Chromium 153 at 1400×900 against vp run dev, main at bb79977 versus this branch (388f12769a at capture time, rebased since with no source change), isolated state with a synthetic project and only Claude enabled. Both builds used real Claude Code 2.1.287 (Agent SDK 0.3.276) with a logged-in account. Each turn was a new thread on Sonnet 5.5 (claude-sonnet-5-5), High effort, with the prompt Reply with exactly: OK; four turns total, one per build and selection, each answered OK in one SDK turn with no tool calls. The env value is read from the query.open options in the native provider log, the window from modelUsage in the SDK result of the same turn.

Selection Reading main this branch
200k settings.env.CLAUDE_CODE_DISABLE_1M_CONTEXT absent "1"
200k modelUsage["claude-sonnet-5-5"].contextWindow 1,000,000 200,000
200k composer meter 19%·38k/200k 16%·31k/200k
1M settings.env.CLAUDE_CODE_DISABLE_1M_CONTEXT absent "0"
1M modelUsage["claude-sonnet-5-5[1m]"].contextWindow 1,000,000 1,000,000
1M composer meter 3.8%·38k/1m 3.8%·38k/1m

The meter and the stored tokenUsage.maxTokens (200,000 and 1,000,000) follow the selected catalog option on both builds, so they do not prove the live window; the native result does. With 1M selected the query model and the result key carry [1m] on both builds, with canonicalModel claude-sonnet-5-5. Each result also lists claude-haiku-4-5-20251001 with a 200,000 window: that is Claude Code's own internal housekeeping call, not a T3 model selection.

Before: 200k selected, Claude Code ran at 1M After: 200k selected, Claude Code ran at 200k
Before: Sonnet 5.5 thread with Context Window 200k selected; the composer meter reads 19%·38k/200k while Claude Code ran the turn with a 1,000,000-token window After: Sonnet 5.5 thread with Context Window 200k selected; the composer meter reads 16%·31k/200k and Claude Code ran the turn with a 200,000-token window

Tests. ClaudeModelCatalog.test.ts: the env for Opus 4.6 by slug and by the opus-4.6 alias at 200k ("1") and 1m ("0"); the catalog default with no selection (Opus 4.6 "0", Sonnet 4.6 "1"); no env for Opus 4.7, Haiku 4.5 and a custom model even with 200k passed. claudeModelOptions.test.ts: the compiled env per window, none for a model without a selector, and a different queryIdentity for 200k and 1m, including on a catalog where both windows share one API model id. ClaudeAdapterV2.test.ts: the SDK query carries settings.env next to fastMode with model claude-opus-4-6 at 200k and claude-opus-4-6[1m] at 1m; the env survives autoCompactWindow; a model without a selector gets no env key; Opus 4.6 usage is projected against 200,000 and 1,000,000 maxTokens. 156 tests pass across the three files; with the three source files reverted, 16 fail (the resolver, compiled env, query identity, SDK env, and the Opus 4.6 meter reporting 1M at 200k). The Claude replay suites (ClaudeReplayFixtures, OrchestratorReplayFixtures, OrchestratorReplayRecovery, OrchestratorReplayRestartBackgroundNote, ThreadFork, ThreadMergeBack, ProviderSwitch) pass against the updated transcripts, 207 tests. Server typecheck, targeted lint, format and knip are clean.

Not checked: models other than Sonnet 5.5 at runtime (the unit tests use Opus 4.6 and Sonnet 4.6; Opus 5, Opus 5.5 and Fable 5 share the same catalog selector); switching 200k and 1m on a live thread (covered by the queryIdentity tests, each runtime turn used a new thread); a CLAUDE_CODE_DISABLE_1M_CONTEXT set in the user's Claude settings or in the server's process environment (the evidence launcher removed inherited overrides); behavior near a full 200k context; remote, tunnel, desktop and mobile clients (same server code).

Implemented with Claude Code (Claude Opus 5.5, coordinated by Claude Fable 5.1); tests and evidence capture by GPT-6 Astra via Codex.

@github-actions github-actions Bot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Oct 3, 2026
macroscopeapp[bot]
macroscopeapp Bot previously approved these changes Oct 3, 2026
@macroscopeapp

macroscopeapp Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The PR changes the effective default Claude context window on existing production paths and makes the server-selected window override user settings for selector models. Despite focused scope and substantial tests, that default behavior change warrants human review.

You can add or adjust custom eligibility rules. Learn more.

@juliusmarminge juliusmarminge added the macroscope-review Opt PRs made by unvouched contributors in for Macroscope review. Vouched contributors auto-reviews label Oct 3, 2026 — with ChatGPT Codex Connector
@macroscopeapp
macroscopeapp Bot dismissed their stale review October 3, 2026 03:27

Dismissing prior approval to re-evaluate 50f96e3

@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: pingdotgg/t3code/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 58d276d9-eae7-4ca2-bedc-59d24ffcae17
📥 Commits

Reviewing files that changed from the base of the PR and between 50f96e3 and 56504a8.

📒 Files selected for processing (41)
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_background_monitor_wake/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_background_subagent_after_root/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_background_subagent_lifecycle/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_background_task_after_root/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_background_task_interrupt/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_background_task_wake/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_background_wake_before_queued_prompt/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_background_wake_before_queued_prompt_no_echo/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_compact_after_peer_turn/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_compact_after_peer_turn_no_echo/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_compact_after_resume_wake/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_idle_resume/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_local_bash_task/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_mcp_tool_presentation/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_nested_background_subagent_wake/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_nested_subagent_model/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_result_is_error/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/claude_subagent_resume_after_restart/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/message_steering/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/multi_turn/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/multi_turn_restart/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/queued_turn/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/simple/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/subagent/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/thread_fork_native/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/thread_fork_native_continue/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/thread_fork_native_fork_local_rollback/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/thread_fork_native_prior_turn/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/thread_fork_native_siblings/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/thread_merge_back_continue/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/thread_merge_back_siblings/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/thread_rollback/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/tool_call_denied_write/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/tool_call_read_only/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/tool_call_read_only_on_request/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/tool_call_restricted_granular/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/tool_call_workspace_never/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/turn_interrupt/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/turn_interrupt_mid_tool/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/turn_interrupt_restart/claude_transcript.ndjson
  • apps/server/src/orchestration-v2/testkit/fixtures/web_search/claude_transcript.ndjson

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 2 remain after this review.


📝 Walkthrough

Walkthrough

Claude context-window selections now resolve to environment settings for 200k and 1m windows. Compiled selections include those settings in query identity. The Claude adapter applies the settings and reports the corresponding token limits.

Changes

Claude context-window selection

Layer / File(s) Summary
Resolve context-window environment
apps/server/src/provider/ClaudeModelCatalog.ts, apps/server/src/provider/ClaudeModelCatalog.test.ts
The catalog resolver maps configured token counts to CLAUDE_CODE_DISABLE_1M_CONTEXT. Tests cover selected and default windows, plus models without a matching selector.
Compile environment and query identity
apps/server/src/claudeModelOptions.ts, apps/server/src/claudeModelOptions.test.ts
Compiled selections now include the resolved environment, and query identity includes that value. Tests cover different context-window selections, including selections with the same API model ID.
Apply selection in Claude adapter
apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.ts, apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.test.ts, apps/server/src/orchestration-v2/testkit/fixtures/*/claude_transcript.ndjson
The adapter applies compiled environment values in SDK settings. The hard-coded 1,000,000-token window now applies only to Opus 4.7. Tests cover settings and reported token limits. Query fixtures now include the context-window environment setting.

Priority: ⬆️ High

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Bug fix · Severity of issue fixed: High

Sequence Diagram(s)

sequenceDiagram
  participant ModelCatalog
  participant SelectionCompiler
  participant ClaudeAdapter
  participant ClaudeCodeSDK
  SelectionCompiler->>ModelCatalog: Resolve selected context-window setting
  ModelCatalog-->>SelectionCompiler: Return environment value
  SelectionCompiler-->>ClaudeAdapter: Return compiled selection and query identity
  ClaudeAdapter->>ClaudeCodeSDK: Apply environment in SDK settings
Loading

Suggested reviewers: juliusmarminge

Merge Risk: ⚪ Minimal · up to 56504

Selected context windows are passed into Claude query settings and distinguish query identity. The reported Sonnet 5.5 checks show the meter following the 200k and 1m selections; no concrete merge blocker is established.

Architecture Summary

Architecture risk: 🔵 Low · up to 56504

The change affects 1 system.

Changed systems: apps/server

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — apps/server (service) was modified; 47 changed files map to changed impact.

Before / after behavior

  • observed — Modified behavior in apps/server/src/claudeModelOptions.test.ts: Imports the synthetic Claude capable-model and catalog fixtures for the added tests.
  • observed — Modified behavior in apps/server/src/claudeModelOptions.test.ts: Adds expectations for the environment values associated with "200k" and "1m", an unset environment when no context selector is supplied, and distinct query identities for those two selections.
  • observed — Modified behavior in apps/server/src/claudeModelOptions.test.ts: Adds a synthetic-catalog test expecting context-window selections to have different query identities even when their API model IDs are equal.
  • observed — Modified behavior in apps/server/src/claudeModelOptions.ts: Imports resolveClaudeCatalogContextWindowEnv for resolving Claude context-window environment settings.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 6 functions across 6 files. (41 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed For [#8405], the reviewed change sets CLAUDE_CODE_DISABLE_1M_CONTEXT from the selected or default window for catalog models with a selector, includes it in queryIdentity, and passes it to the SDK …
Out of Scope Changes check ✅ Passed The changes remain within [#8405]'s Claude server and adapter scope. The Opus 4.6 meter correction follows its selected context window. The replay fixture updates record the new query environment for …
Title check ✅ Passed The title clearly identifies the server fix that makes the Claude 200k context-window selection take effect.
Description check ✅ Passed The description covers the problem, changes, scope and approval, and verification. It also reports runtime limitations and untested cases.
Full details: Docstring Coverage

Explanation

Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 6 functions across 6 files. (41 skipped: 41 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@Mnigos
Mnigos force-pushed the claude-200k-context-window branch from 50f96e3 to 56504a8 Compare October 3, 2026 03:35
@github-actions github-actions Bot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Oct 3, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

macroscope-review Opt PRs made by unvouched contributors in for Macroscope review. Vouched contributors auto-reviews size:L 100-499 changed lines (additions + deletions). vouch:unvouched PR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Claude 200k context window selection is a no-op; sessions always run at 1M

2 participants