Conversation
Keep Transformers initialization consistent with the plugin flag without changing the training configuration, and restore the state when model loading fails. Signed-off-by: 1fanwang <1fannnw@gmail.com>
Exercise dense model loading and inference inside the disabled context, then confirm ordinary ZeRO-3 partitioning still resumes afterward. Signed-off-by: 1fanwang <1fannnw@gmail.com>
1fanwang
marked this pull request as ready for review
September 23, 2026 16:05
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Disabling ZeRO-3 initialization leaves Transformers' partitioning check enabled. The plugin reports
False, but Transformers reportsTrue, so model loading can still enter the ZeRO-3 path. An exception inside the context also leaves the plugin flag changed.The initialization-only configuration now respects the flag without changing the engine's training configuration. Context exit restores the previous state, including after an exception.
Related to #3170. A two-H100 regression reproduces the reported embedding error and passes with this fix; the original eight-GPU diffusion workflow has not been rerun.
Testing Done
On two H100s connected by NVLink, the same regression was run with the baseline package and the fixed package. The baseline run retained the new test files. The stack was CUDA 12.9, PyTorch 2.11.0+cu129, Transformers 5.15.0 and DeepSpeed 0.18.8; both runs used the same synthetic BERT configuration without downloading a model.
Before, both ranks failed with exit 1:
After, both ranks passed with exit 0. The outside-context control confirms that normal ZeRO-3 partitioning still works:
Initialization and exception-restoration probe
I ran the same probe on the baseline and this branch with Python 3.12.13, PyTorch 2.14.0, Transformers 5.17.0 and DeepSpeed 0.19.7 on macOS arm64.
Save the probe as
probe_zero3_context.pyand run it with the selected checkout installed:Before, exit 1:
After, exit 0:
This also verifies restoration after a caller raises inside the context.
Before submitting
Pull Request section?
to it if that's the case.
documentation guidelines, and
here are tips on formatting docstrings.
Who can review?
Anyone in the community is free to review the PR once the tests have passed. Feel free to tag
members/contributors who may be interested in your PR.