Wait for coordinator to finish recovery before connecting in gpstart - #1898
Merged
Conversation
Contributor
Author
|
To trigger issue you need to dirty your buffers and run immediate stop: and then gpstop -a -i && gpstart -a |
In CBDB3 after the upstream PG16 rebase on commit 7ff23c6 PMSIGNAL_RECOVERY_STARTED is now sent during crash recovery too. With hot_standby=off (the default), this causes postmaster to write PM_STATUS_STANDBY to the pidfile as soon as recovery begins, and pg_ctl -w treats standby as success. gpstart then hits FATAL "Hot standby mode is disabled". Add a shared _waitForCoordinatorRecovery helper that polls dbconn.connect with a 5s interval up to 300s, retrying only on recovery-related FATAL messages (not accepting connections / not yet accepting connections /
leborchuk
approved these changes
Aug 18, 2026
leborchuk
left a comment
Contributor
There was a problem hiding this comment.
LGTM, 300 seconds for recovery wait is normal. And also for large installations we should have option to increase it
Contributor
Author
|
Another option is to change pg_ctl wait semantic for CBDB. But I consider this as undesirable discrepancy with upstream. |
13 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
In CBDB3 after the upstream PG16 rebase on commit 7ff23c6
PMSIGNAL_RECOVERY_STARTED is now sent during crash recovery too.
With hot_standby=off (the default), this causes postmaster to write PM_STATUS_STANDBY to the pidfile as soon as recovery begins, and pg_ctl -w treats standby as success. gpstart then hits
FATAL "Hot standby mode is disabled".
Add a shared _waitForCoordinatorRecovery helper that polls dbconn.connect with a 5s interval up to 300s, retrying only on recovery-related FATAL messages (not accepting connections / not yet accepting connections /