docs(streaming): clarify warm_willneed call paths - #120
Open
agourakis82 wants to merge 4 commits into
Open
agourakis82 wants to merge 4 commits into
agourakis82 wants to merge 4 commits into
Conversation
While investigating Edge0-AI#110 (edge0-8b first-token latency), traced why toggling warm_willneed had no visible effect on that tier: it's only read by StreamingSwitchGLU.prefetch() and stage_experts(), and stage_experts() early-returns immediately when staged is False. prod_k8() (edge0-8b's production profile) sets staged=False and history_prefetch=False, so neither call site that reads warm_willneed ever executes -- the flag is dead code for that tier's default config, not merely ineffective in this instance. Documents this in both the class docstring (options.py) and the streaming.md options table, so the next person investigating slow first-token latency on edge0-8b doesn't spend time on the same dead end. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Preserve the PR history by merging current upstream main. Adapt the documentation to the python/ layout and describe explicit prefetch and current staged consumer behavior without changing runtime options. Assisted-by: Codex
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The documentation-only changes have one minor clarification remaining and no blocking issues.
Review effort: Balanced
Findings: 1
What changed in this PR
Clarifies when streaming readahead applies, helping users investigating first-token latency in #110.
Changes:
- Documents
warm_willneedcall paths and loading limitations. - Updates the
prod_k8()description to reflect staged prerouter consumer layers.
| File | Description |
|---|---|
| python/src/edge0/streaming/options.py | Expands the warm_willneed field documentation. |
| docs/streaming.md | Clarifies readahead behavior and the production preset. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+95
to
+97
| #: or affect plain on-demand / whole-layer loading. Explicit prefetch | ||
| #: calls, including ``prefetch_from_prefill()``, can still consult it | ||
| #: when automatic history prefetch and staging are disabled. |
Contributor
Author
There was a problem hiding this comment.
Fixed both passages: the direct example is now prefetch(experts), and prefetch_from_prefill() explicitly requires an expert set captured while staging was enabled. The text states that the helper is a no-op when staging is disabled from initialization. Verified the call paths locally; no runtime changes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Clarifies the
warm_willneedcall paths discussed in #110.prefetch(experts)can issue readahead for missing experts even when history prefetch and staging are disabled.prefetch_from_prefill()needs an expert set captured while staging was enabled. With staging disabled from initialization, it is a no-op.prod_k8()now documents staged decode for prerouter consumer layers.Addresses the inline review and merges current
mainwithout rewriting the PR history. Relative tomain, onlydocs/streaming.mdand comments inpython/src/edge0/streaming/options.pychange; runtime code and settings are unchanged.Validation on macOS with MLX 0.30.6: 128 repository tests and six local call-path/AST checks passed; four slow tests were deselected. The AST check confirms executable options match
main.git diff --checkpassed. This documentation update does not claim a new performance measurement or a fix for first-token latency.