Conversation
ApprovabilityVerdict: Not approved Macroscope's review found this PR not approvable — The PR changes the effective default Claude context window on existing production paths and makes the server-selected window override user settings for selector models. Despite focused scope and substantial tests, that default behavior change warrants human review. You can add or adjust custom eligibility rules. Learn more. |
Dismissing prior approval to re-evaluate 50f96e3
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (41)
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 2 remain after this review. 📝 WalkthroughWalkthroughClaude context-window selections now resolve to environment settings for 200k and 1m windows. Compiled selections include those settings in query identity. The Claude adapter applies the settings and reports the corresponding token limits. ChangesClaude context-window selection
Priority: ⬆️ High Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Bug fix · Severity of issue fixed: High Sequence Diagram(s)sequenceDiagram
participant ModelCatalog
participant SelectionCompiler
participant ClaudeAdapter
participant ClaudeCodeSDK
SelectionCompiler->>ModelCatalog: Resolve selected context-window setting
ModelCatalog-->>SelectionCompiler: Return environment value
SelectionCompiler-->>ClaudeAdapter: Return compiled selection and query identity
ClaudeAdapter->>ClaudeCodeSDK: Apply environment in SDK settings
Suggested reviewers: Merge Risk: ⚪ Minimal · up to Selected context windows are passed into Claude query settings and distinguish query identity. The reported Sonnet 5.5 checks show the meter following the 200k and 1m selections; no concrete merge blocker is established. Architecture SummaryArchitecture risk: 🔵 Low · up to The change affects 1 system. Changed systems: Architecture concerns Review detailsSystems and components
Before / after behavior
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 6 functions across 6 files. (41 skipped: 41 unsupported.)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
50f96e3 to
56504a8
Compare
Fixes #8405.
Problem
Claude Code runs a 1M-capable model at 1M when given its bare slug, unless
CLAUDE_CODE_DISABLE_1M_CONTEXTis set. T3 encoded the context window only in the slug ([1m]for 1M, the bare slug for 200k), so selecting 200k did nothing and the session still ran at 1,000,000 tokens. Onmainthe composer meter and the storedtokenUsage.maxTokenscome from the selected option, so they read 200k while Claude Code reported a 1,000,000-token window for the same turn. Separately,claudeContextWindowhardcoded 1M forclaude-opus-4-6, which has a 200k/1m selector, so its meter showed 1M with 200k selected.Change
ClaudeModelCatalog.ts: newresolveClaudeCatalogContextWindowEnv. For a model whose catalog entry has a context window selector, it returnsCLAUDE_CODE_DISABLE_1M_CONTEXT: "1"when the selected (or default) window is 200k or less and"0"otherwise. The window is stated both ways so the selection, not anenvin the user's Claude settings, decides it; Claude Code also refuses a[1m]slug while the variable is set. Models without a selector (Opus 4.7, Haiku 4.5, custom models) get no value and keep whatever the user configured.claudeModelOptions.ts:compileClaudeModelSelectionreturns thatenvand includes it inqueryIdentity. The adapter reuses a live query only when the identity matches, and Claude Code reads the env when it starts, so switching between 200k and 1m on a thread opens a new query, also for a catalog model whose API id does not change with the window.ClaudeAdapterV2.ts:makeClaudeQueryOptionspasses the env assettings.envon the SDK query, next to the selection's other settings. These are flag settings, which outrankenvfrom the user's settings files. The existing merge with the instance's SDK settings andautoCompactWindowis unchanged and keeps the env.ClaudeAdapterV2.ts:claudeContextWindowno longer hardcodes 1M forclaude-opus-4-6. Opus 4.6 has a 200k/1m selector in the catalog, so its storedmaxTokensand meter now follow the selection like the other selector models.claude-opus-4-7has no selector and stays at a fixed 1M.Scope and approval
Triaged bug #8405, following the triage comment, which adds the acceptance case from #5286 (start a new Claude session, select 200k, read the meter): "After the provider fix, this sequence must make both the live provider window and the meter show 200k." Server and Claude adapter only; no contract or client change, since the clients already render the stored
maxTokens. The recorded Claude replay transcripts (41 files underorchestration-v2/testkit/fixtures) now expectsettings.env.CLAUDE_CODE_DISABLE_1M_CONTEXT: "1"on their 59query.openframes, because they were recorded onclaude-sonnet-4-6, whose catalog default is 200k.One behavior change for selector models: with no explicit selection, the catalog default is now stated (Opus 4.6 defaults to 1M and gets
"0", Sonnet 4.6 defaults to 200k and gets"1"), so aCLAUDE_CODE_DISABLE_1M_CONTEXTin a user's Claude settings no longer decides the window for these models; the composer selection does. Models whose default is 200k (Sonnet 5.5, Sonnet 5, Sonnet 4.6) used to run at 1M on a new thread while the meter showed 200k; they now run at the 200k the meter shows, and 1M is an explicit selection.Not changed:
ClaudeTextGeneration(titles, commit messages) still passes only the model id to its one-shotclaude -pcalls.Verification
Observed result. Real web client, headless Playwright Chromium 153 at 1400×900 against
vp run dev,mainat bb79977 versus this branch (388f12769a at capture time, rebased since with no source change), isolated state with a synthetic project and only Claude enabled. Both builds used real Claude Code 2.1.287 (Agent SDK 0.3.276) with a logged-in account. Each turn was a new thread on Sonnet 5.5 (claude-sonnet-5-5), High effort, with the promptReply with exactly: OK; four turns total, one per build and selection, each answeredOKin one SDK turn with no tool calls. The env value is read from thequery.openoptions in the native provider log, the window frommodelUsagein the SDK result of the same turn.mainsettings.env.CLAUDE_CODE_DISABLE_1M_CONTEXT"1"modelUsage["claude-sonnet-5-5"].contextWindow19%·38k/200k16%·31k/200ksettings.env.CLAUDE_CODE_DISABLE_1M_CONTEXT"0"modelUsage["claude-sonnet-5-5[1m]"].contextWindow3.8%·38k/1m3.8%·38k/1mThe meter and the stored
tokenUsage.maxTokens(200,000 and 1,000,000) follow the selected catalog option on both builds, so they do not prove the live window; the native result does. With 1M selected the query model and the result key carry[1m]on both builds, withcanonicalModelclaude-sonnet-5-5. Each result also listsclaude-haiku-4-5-20251001with a 200,000 window: that is Claude Code's own internal housekeeping call, not a T3 model selection.Tests.
ClaudeModelCatalog.test.ts: the env for Opus 4.6 by slug and by theopus-4.6alias at 200k ("1") and 1m ("0"); the catalog default with no selection (Opus 4.6"0", Sonnet 4.6"1"); no env for Opus 4.7, Haiku 4.5 and a custom model even with 200k passed.claudeModelOptions.test.ts: the compiled env per window, none for a model without a selector, and a differentqueryIdentityfor 200k and 1m, including on a catalog where both windows share one API model id.ClaudeAdapterV2.test.ts: the SDK query carriessettings.envnext tofastModewith modelclaude-opus-4-6at 200k andclaude-opus-4-6[1m]at 1m; the env survivesautoCompactWindow; a model without a selector gets noenvkey; Opus 4.6 usage is projected against 200,000 and 1,000,000maxTokens. 156 tests pass across the three files; with the three source files reverted, 16 fail (the resolver, compiled env, query identity, SDK env, and the Opus 4.6 meter reporting 1M at 200k). The Claude replay suites (ClaudeReplayFixtures,OrchestratorReplayFixtures,OrchestratorReplayRecovery,OrchestratorReplayRestartBackgroundNote,ThreadFork,ThreadMergeBack,ProviderSwitch) pass against the updated transcripts, 207 tests. Server typecheck, targeted lint, format and knip are clean.Not checked: models other than Sonnet 5.5 at runtime (the unit tests use Opus 4.6 and Sonnet 4.6; Opus 5, Opus 5.5 and Fable 5 share the same catalog selector); switching 200k and 1m on a live thread (covered by the
queryIdentitytests, each runtime turn used a new thread); aCLAUDE_CODE_DISABLE_1M_CONTEXTset in the user's Claude settings or in the server's process environment (the evidence launcher removed inherited overrides); behavior near a full 200k context; remote, tunnel, desktop and mobile clients (same server code).Implemented with Claude Code (Claude Opus 5.5, coordinated by Claude Fable 5.1); tests and evidence capture by GPT-6 Astra via Codex.