Repository navigation
Conversation
ghwang-s
marked this pull request as ready for review
October 5, 2026 12:33
Michael-RDev
reviewed
Oct 6, 2026
Michael-RDev
left a comment
There was a problem hiding this comment.
The implementation looks good to me. Waiting until DeepSpeed finishes setting up the optimizer makes sense here, since it preserves each group鈥檚 lr when warmup_max_lr is "auto" and it handles the Adagrad optimizer replacement.
- One nonblocking suggestion: could a small version of the real-engine Muon/Adam checkpoint reproduction described in the PR be added to the automated tests? The current tests mock engine initialization and restore state into the same objects. Creating a fresh engine, loading the checkpoint, and comparing the next step鈥檚 learning rates and parameters against uninterrupted training would help protect this initialization path.
All 14 new tests and the independent single-rank CPU/Gloo resume checks passed locally. CUDA and multi-rank behavior weren鈥檛 independently verified in this review, but everything looks good to me.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Following up on deepspeedai/DeepSpeed#7713. With Muon at
0.02and Adam at1e-5, settingwarmup_max_lr: "auto"currently gives both groups1e-5. We fill this value from the dummy optimizer's single LR before DeepSpeed creates the actual parameter groups.This PR resolves
autofrom those groups afterdeepspeed.initialize. The values are passed as a list, so this also works with older DeepSpeed versions that read only the first group when givenNone. Explicit LR settings and custom scheduler callbacks keep their existing behavior.The timing matters for Adagrad: ZeRO can rebuild the optimizer during initialization, while DeepSpeed's scheduler callback receives the original
basic_optimizer. Creating the scheduler afterward attaches it to the optimizer used for training.All 14 CPU regression tests pass with DeepSpeed 0.19.2 and current main;
make qualityalso passes. I also ran real two-rank CPU/Gloo engines with ZeRO-2 for both Muon/Adam and Adagrad. Each completed 12 updates and resumed in fresh processes with matching parameters, losses, and LR histories. The Muon/Adam peaks stay at[0.02, 1e-5].The attached reproduction includes the commands, CPU launch setup, and recorded results. It also records a ZeRO-3 Adagrad checkpoint failure reproduced with an explicit-LR control. CUDA validation is still pending.
accelerate4367-engine-repro.zip