Skip to content

Fix DeepSpeed auto warmup LR for multiple optimizer groups - #4367

Open
ghwang-s wants to merge 5 commits into
huggingface:mainfrom
ghwang-s:codex/fix-deepspeed-auto-group-lr
Open

ghwang-s wants to merge 5 commits into
huggingface:mainfrom
ghwang-s:codex/fix-deepspeed-auto-group-lr

Conversation

@ghwang-s

@ghwang-s ghwang-s commented Oct 5, 2026 •

Copy link
Copy Markdown

Following up on deepspeedai/DeepSpeed#7713. With Muon at 0.02 and Adam at 1e-5, setting warmup_max_lr: "auto" currently gives both groups 1e-5. We fill this value from the dummy optimizer's single LR before DeepSpeed creates the actual parameter groups.

This PR resolves auto from those groups after deepspeed.initialize. The values are passed as a list, so this also works with older DeepSpeed versions that read only the first group when given None. Explicit LR settings and custom scheduler callbacks keep their existing behavior.

The timing matters for Adagrad: ZeRO can rebuild the optimizer during initialization, while DeepSpeed's scheduler callback receives the original basic_optimizer. Creating the scheduler afterward attaches it to the optimizer used for training.

All 14 CPU regression tests pass with DeepSpeed 0.19.2 and current main; make quality also passes. I also ran real two-rank CPU/Gloo engines with ZeRO-2 for both Muon/Adam and Adagrad. Each completed 12 updates and resumed in fresh processes with matching parameters, losses, and LR histories. The Muon/Adam peaks stay at [0.02, 1e-5].

The attached reproduction includes the commands, CPU launch setup, and recorded results. It also records a ZeRO-3 Adagrad checkpoint failure reproduced with an explicit-LR control. CUDA validation is still pending.

accelerate4367-engine-repro.zip

@ghwang-s
ghwang-s marked this pull request as ready for review October 5, 2026 12:33

@Michael-RDev Michael-RDev left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implementation looks good to me. Waiting until DeepSpeed finishes setting up the optimizer makes sense here, since it preserves each group鈥檚 lr when warmup_max_lr is "auto" and it handles the Adagrad optimizer replacement.

  • One nonblocking suggestion: could a small version of the real-engine Muon/Adam checkpoint reproduction described in the PR be added to the automated tests? The current tests mock engine initialization and restore state into the same objects. Creating a fresh engine, loading the checkpoint, and comparing the next step鈥檚 learning rates and parameters against uninterrupted training would help protect this initialization path.

All 14 new tests and the independent single-rank CPU/Gloo resume checks passed locally. CUDA and multi-rank behavior weren鈥檛 independently verified in this review, but everything looks good to me.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants