Skip to content

Fix Claude context usage for custom models: send the 1M the picker shows, and measure the streamed usage - #800

Open
AndPuQing wants to merge 2 commits into
zeronsh:mainfrom
AndPuQing:zeron/zeron-cc-context-windows-cu
Open

AndPuQing wants to merge 2 commits into
zeronsh:mainfrom
AndPuQing:zeron/zeron-cc-context-windows-cu

Conversation

@AndPuQing

@AndPuQing AndPuQing commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Symptom

A custom gateway model (deepseek-v4.1-flash, configured through ANTHROPIC_MODEL) shows the context ring pinned at 0%, and a model whose catalog row advertises a 1M window actually runs at 200K. Three defects sit behind those two numbers; the ring only reads true when all three are fixed, which is why they ship together.

1. The Run dropped the context-window choice the picker displayed

Claude Code selects the 1M window with a model-id suffix (<id>[1m]). The catalog advertises the orphan row deepseek-v4.1-flash[1m]; the picker folds it into deepseek-v4.1-flash + a contextWindow option whose default is 1m, and renders that default as the selected value. But the Run carried only explicit picks, so model_options stayed empty and build_command (crates/harness/src/claude/mod.rs:193) — which appends [1m] only when contextWindow == "1m" — sent the bare id. Verified against CLI 2.1.289 on the reporting gateway:

--model deepseek-v4.1-flash        → modelUsage[…].contextWindow = 200000
--model "deepseek-v4.1-flash[1m]"  → modelUsage[…].contextWindow = 1000000

So the ring's capacity was 5× too small, and everything the CLI derives from the window — auto-compact in particular — was keyed to 200K.

Fix: the composer's Run payload materializes each offered option's default (materialized_options + Pickers::run_model_options). Chat rows still store explicit picks only, so no catalog default is frozen into a chat.

2. The aggregated assistant frame reports all-zero usage → 0% (upstream #786)

The ring measures the latest parent assistant message's input + cache_read + cache_creation. On this gateway the aggregated assistant frame reports all zeros for those counters while message_delta carries the real numbers:

stream_event message_start  usage={input:0, output:0, cache_creation:0, cache_read:0}
assistant                   usage={input:0, output:0, cache_creation:0, cache_read:0}
stream_event message_delta  usage={input:9515, output:2, cache_creation:0, cache_read:2048}
result modelUsage["deepseek-v4.1-flash"].contextWindow = 200000

The normalizer dropped every stream_event except content_block_delta, so the only real measurement never reached the meter; the frame's Some(0) then won the partial merge (tokens.or(previous), crates/doc/src/schema.rs:360) and pinned the ring at 0% for the rest of the session. Note the assistant frame arrives before message_delta on this wire (line 6 vs line 8 of the probe), so the streamed frames must emit their own measurement rather than act as a fallback for the frame.

Fix: message_start/message_delta usage is a measurement source (parent-only, positive-only), and an all-zero aggregated frame no longer publishes a zero that would erase a real measurement.

3. The result-frame window lookup missed the [1m] entry

With --model <id>[1m] the CLI reports the assistant frame's model without the suffix but keeps it in modelUsage:

assistant  model = "deepseek-v4.1-flash"
result     modelUsage["deepseek-v4.1-flash[1m]"].contextWindow = 1000000

The lookup matched only the bare id / canonicalModel, so the single-entry fallback masked this in simple sessions — but a turn that also used a subagent (2+ entries) silently lost the capacity ("—"). Found while validating (1); fixed by trying both spellings before the fallback.

Relation to upstream work

Verification

  • Live probes against Claude Code 2.1.289 on the reporting gateway (frames quoted above; the probe used the exact flags build_command emits).
  • cargo test -p zeron-harness (all targets) and cargo test -p zeron-ui --lib green. A few UI layout/lightbox tests fail under parallel load on this machine; they are pre-existing flakes and pass in isolation.
  • New tests: the exact [Bug]: Context-window ring always shows 0% for Claude Code (cache_read_input_tokens not counted) #786 frame sequence yields one ContextUsage with the real sum and never a zero; subagent stream frames never reach the parent meter; [1m] / canonicalModel modelUsage matching with a second model present; option defaults materialize into the Run while resolved() and chat rows stay default-free.

Known follow-ups (not in this PR)

The mobile client (crates/client/src/session/mod.rs:1051) and the engine's restart+steer rebuild (request_from_chat_row) send only the chat row's persisted options, so a folded 1M row keeps the pre-fix behavior on those two paths; pinning it would mean writing the defaults into the chat config at create time.

How to check in the app

Open a chat on the custom model: the ring should read ≈1% with the popover showing 11,563 / 1,000,000 tokens, instead of 0% against a 200K capacity. A turn that also runs a subagent keeps the capacity (previously it fell back to "—").

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Three defects made the context ring wrong for a model whose catalog row
advertises a 1M window:

- The Run carried only explicit option picks, so a folded `<model>[1m]`
  row (displayed with the 1M Context Window trait) ran as the bare id and
  the CLI used its 200K window. The Run now materializes each offered
  option's default; chat rows still persist explicit picks only.
- Gateways that zero the input counters on the aggregated `assistant`
  frame pinned the ring at 0%: the normalizer dropped every stream_event
  except deltas, and the frame's Some(0) won the partial merge. Streamed
  `message_start`/`message_delta` usage is now a measurement source of its
  own (parent-only, positive-only), and an all-zero frame no longer
  publishes a zero that would erase a real measurement.
- With the `[1m]` suffix the CLI strips the suffix from the assistant
  frame's model but keeps it in `modelUsage`, so the result-frame window
  lookup missed the entry. Both spellings (key and canonicalModel) are
  tried before the single-entry fallback.

Fixes upstream zeronsh#786.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant