fix(server): retry all PD peers while waiting for storage - #3129
Conversation
There was a problem hiding this comment.
Pull request overview
This pull request fixes wait-storage.sh behavior in multi-PD HStore deployments by avoiding “pinning” to the first PD that answers /v1/health, and instead probing store readiness (/v1/stores) across all configured PD REST peers until any peer reports an Up store. This improves Server startup robustness when some PD peers are listening but not yet raft-ready.
Changes:
- Update
wait-storage.shto retry/v1/storesacross every configured PD peer on each retry cycle and succeed as soon as any peer reports"state":"Up". - Add a deterministic shell regression test suite that mocks
curl,sleep, andtimeoutto validate failover / retry behavior. - Add a dedicated GitHub Actions job to run the new shell regression suite in CI.
Reviewed changes
Copilot reviewed 2 out of 3 changed files in this pull request and generated no comments.
| File | Description |
|---|---|
hugegraph-server/hugegraph-dist/src/assembly/static/bin/wait-storage.sh |
Removes /v1/health gating and probes /v1/stores across all PD peers per retry, succeeding on the first peer that reports an Up store. |
hugegraph-server/hugegraph-dist/src/assembly/travis/test-wait-storage.sh |
Adds deterministic, mocked shell tests covering peer order, retry/failover, and timeout behavior. |
.github/workflows/server-ci.yml |
Adds a CI job that runs the new test-wait-storage.sh regression suite. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
- cap each PD connection attempt at 2 seconds - cap each PD request at 3 seconds - cover failover after a hanging first peer
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #3129 +/- ##
============================================
- Coverage 39.19% 38.20% -0.99%
- Complexity 264 424 +160
============================================
Files 770 770
Lines 65779 65779
Branches 8726 8726
============================================
- Hits 25779 25131 -648
- Misses 37247 37927 +680
+ Partials 2753 2721 -32 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
imbajin
left a comment
There was a problem hiding this comment.
TODO: ensure the image version → ubuntu22?
Purpose of the PR
wait-storage.shselected the first PD answering/v1/health, then polled/v1/storesonly on that peer. Since the health endpoint can succeed beforethe peer has raft-backed store state, a storeless first peer could block Server
startup even when another configured PD already reported an
Upstore.Main Changes
/v1/storesacross every configured PD peer on every retry.Up, without a separate/v1/healthselection stage.timing, and the existing outer timeout failure.
Verifying these changes
test-wait-storage.sh: 4 scenarios passed, 0 failed.bash -npassed for the production and test scripts.git diff --checkpassed.Does this PR potentially affect the following parts?
Documentation Status
Doc - TODODoc - DoneDoc - No NeedOut of scope
wait-partition.sh