Skip to content

docs(streaming): clarify warm_willneed call paths - #120

Open
agourakis82 wants to merge 4 commits into
Edge0-AI:mainfrom
agourakis82:docs/warm-willneed-inert-under-prod-k8
Open

agourakis82 wants to merge 4 commits into
Edge0-AI:mainfrom
agourakis82:docs/warm-willneed-inert-under-prod-k8

Conversation

@agourakis82

@agourakis82 agourakis82 commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Clarifies the warm_willneed call paths discussed in #110.

  • Direct prefetch(experts) can issue readahead for missing experts even when history prefetch and staging are disabled.
  • prefetch_from_prefill() needs an expert set captured while staging was enabled. With staging disabled from initialization, it is a no-op.
  • The flag does not enable prefetch/staging or warm plain on-demand and whole-layer loads by itself.
  • prod_k8() now documents staged decode for prerouter consumer layers.

Addresses the inline review and merges current main without rewriting the PR history. Relative to main, only docs/streaming.md and comments in python/src/edge0/streaming/options.py change; runtime code and settings are unchanged.

Validation on macOS with MLX 0.30.6: 128 repository tests and six local call-path/AST checks passed; four slow tests were deselected. The AST check confirms executable options match main. git diff --check passed. This documentation update does not claim a new performance measurement or a fix for first-token latency.

agourakis82 and others added 2 commits September 22, 2026 21:11
While investigating Edge0-AI#110 (edge0-8b first-token latency), traced why
toggling warm_willneed had no visible effect on that tier: it's only
read by StreamingSwitchGLU.prefetch() and stage_experts(), and
stage_experts() early-returns immediately when staged is False.
prod_k8() (edge0-8b's production profile) sets staged=False and
history_prefetch=False, so neither call site that reads warm_willneed
ever executes -- the flag is dead code for that tier's default config,
not merely ineffective in this instance.

Documents this in both the class docstring (options.py) and the
streaming.md options table, so the next person investigating slow
first-token latency on edge0-8b doesn't spend time on the same dead
end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Preserve the PR history by merging current upstream main. Adapt the documentation to the python/ layout and describe explicit prefetch and current staged consumer behavior without changing runtime options.

Assisted-by: Codex
Copilot AI balanced review requested due to automatic review settings October 3, 2026 18:18
@agourakis82 agourakis82 changed the title docs(streaming): document warm_willneed as inert under prod_k8() docs(streaming): clarify warm_willneed call paths Oct 3, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The documentation-only changes have one minor clarification remaining and no blocking issues.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
What changed in this PR

Clarifies when streaming readahead applies, helping users investigating first-token latency in #110.

Changes:

  • Documents warm_willneed call paths and loading limitations.
  • Updates the prod_k8() description to reflect staged prerouter consumer layers.
File Description
python/​src/​edge0/​streaming/​options.py Expands the warm_willneed field documentation.
docs/​streaming.md Clarifies readahead behavior and the production preset.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread python/src/edge0/streaming/options.py Outdated
Comment on lines +95 to +97
#: or affect plain on-demand / whole-layer loading. Explicit prefetch
#: calls, including ``prefetch_from_prefill()``, can still consult it
#: when automatic history prefetch and staging are disabled.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed both passages: the direct example is now prefetch(experts), and prefetch_from_prefill() explicitly requires an expert set captured while staging was enabled. The text states that the helper is a no-op when staging is disabled from initialization. Verified the call paths locally; no runtime changes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants