Repository navigation
Conversation
`load_accelerator_state` reported a failure to restore the random states with `logger.info`, which is below the effective log level of a default setup (`logging` inherits `WARNING` from the root logger). A checkpoint whose RNG state is missing or unreadable therefore resumed silently, using different random numbers than the interrupted run. Log it at `WARNING` instead, on every process since each rank restores its own file, and include the checkpoint path and the underlying error. Resuming from a checkpoint without RNG state still works. Fixes huggingface#4283 Signed-off-by: CJstate <142857225+CJstate@users.noreply.github.com>
CJstate
added a commit
to CJstate/CJstate
that referenced
this pull request
Oct 2, 2026
Author
|
Small process note on CI: the five Could someone click "Approve and run workflows" when convenient? The description has the local before/after evidence (before: effective level 30, zero output; after: exactly one |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Fixes #4283.
load_accelerator_state()restores the RNG state inside atry/except, and reported a failure withlogger.info("Could not load random states"). That message is below the effective log level of a default setup:accelerate.checkpointinghas no level of its own, sologginginheritsWARNINGfrom the root logger and the record is dropped before it reaches any handler. A checkpoint whose RNG file is missing, truncated, or unreadable therefore resumed with no output at all, silently using different random numbers (and astepof 0, sincestepis stored in the same file) than the interrupted run.This PR reports the failure at
WARNINGinstead, on every process (each rank restores its own file), and includes the checkpoint path and the underlying error:Resuming from a checkpoint without RNG state still works, it is just not hidden anymore.
Before / after, on the reproduction from #4283
Using the script attached to the issue, unmodified:
# after RNG restored: False RNG diverged: True override_attributes: {} effective log level: 30 LOG[WARNING] accelerate.checkpointing: [RANK 0] Could not load the random states from .../random_states_0.pkl: Weights only load failed. ... Training will resume, but the random number generators were not restored, so this run may not reproduce the interrupted one.Note about prior art: #4308 reached the same conclusion (this is the same fix, expressed independently). It was closed without review on 2026-09-25, in the same batch of PRs as #4285, #4239 and ~13 others, and the failure is still present on
main. Since #4283 is still open, this PR restores the fix.Tests
Three regression tests in
tests/test_state_checkpointing.py::RandomStateRestoreTest:test_unreadable_rng_state_warns: save a checkpoint, corruptrandom_states_0.pkl, resume → exactly oneWARNINGnaming the file, and resuming still returns.test_missing_rng_state_warns: same for an absent RNG file.test_readable_rng_state_restores_without_warning: a healthy checkpoint loads with no warning, restores Python/NumPy/PyTorch RNG state exactly, and restoresstep.Verified on Windows 11 / Python 3.13 / PyTorch 2.9.1:
git stashofsrc/accelerate/checkpointing.py):2 failed, 1 passed— both failures areAssertionError: no logs of level WARNING or higher triggered on accelerate.checkpointing.3 passed.21 passed(no other test relies on the INFO message).ruff checkandruff format --check(0.13.1, the version pinned insetup.py) pass on both files.Before submitting
to it if that's the case. → [Bug] RNG state restoration failure swallowed silently — resumed training diverges without warning #4283
(
random_states_*.pklonly appears in the FSDP usage guide as part of a checkpoint listing), so there is nothingto update.
Who can review?
@BenjaminBossan @SunMarc (core parts of the library / checkpointing)