Repository navigation
Apply configured store retry budget - #90
behinddwalls wants to merge 1 commit into
Conversation
|
Careful with this one. The SDK retry setting applies to every call, including the manifest CAS. If a CAS lands but its reply gets lost, the SDK retries, gets a 412 from our own write, and publish.rs takes that as a lost race and deletes the log segment the new manifest points at. I reproduced it with a fault that applies the write and then answers 412, and main already hits this with the SDK default of 3 attempts, so more retries make it more likely. Could the retries stay on reads, with conditional writes and deletes at one attempt? I opened #103 for the publisher side. |
## Summary ### Why? Backend SDK retries covered conditional mutations as well as reads. If a manifest CAS landed but its response was lost, an SDK retry could receive 412 from the already-committed write and make the publisher treat success as a lost race. ### What? Apply `store.max_retries` only to idempotent GCS and S3 reads and interrupted bulk reads. Configure GCS mutation clients and the S3 mutation client for one attempt, while S3 metadata reads use a separate retry-enabled client and presigned GETs retain explicit bounded retries. Healthy calls remain one request. Failure paths add at most `store.max_retries` read attempts; conditional writes, uploads, copies, multipart operations, and deletes remain single-attempt. ## Test Plan ✅ `cargo check -p walgit-store --all-features` ✅ `cargo test -p walgit-store --all-features` ✅ `cargo test -p walgit-server --test sim healthy_request_round_trip_budgets -- --exact` ✅ `cargo clippy -p walgit-store --all-targets --all-features -- -D warnings` ✅ `cargo fmt --all -- --check` ✅ `git diff --check` ## Issue Closes tobi#82
a46bb4b to
34f240d
Compare
|
Restricted configured retries to idempotent reads in [addressed by agent] |
|
Thanks, this fixes the risky part, conditional writes aren't retried anymore. I think it went a bit further than it needs to though. On GCS everything is single-attempt now, including normal reads and large uploads, and on S3 the multipart part uploads too. Those retries were safe and worth keeping, it's only the conditional writes that need to be one try. Small one: the comment at s3.rs:313 still says the SDK retries three times. |
Summary
Why?
Backend SDK retries covered conditional mutations as well as reads. If a manifest CAS landed but its response was lost, an SDK retry could receive 412 from the already-committed write and make the publisher treat success as a lost race.
What?
Apply
store.max_retriesonly to idempotent GCS and S3 reads and interrupted bulk reads. Configure GCS mutation clients and the S3 mutation client for one attempt, while S3 metadata reads use a separate retry-enabled client and presigned GETs retain explicit bounded retries.Healthy calls remain one request. Failure paths add at most
store.max_retriesread attempts; conditional writes, uploads, copies, multipart operations, and deletes remain single-attempt.Test Plan
✅
cargo check -p walgit-store --all-features✅
cargo test -p walgit-store --all-features✅
cargo test -p walgit-server --test sim healthy_request_round_trip_budgets -- --exact✅
cargo clippy -p walgit-store --all-targets --all-features -- -D warnings✅
cargo fmt --all -- --check✅
git diff --checkIssue
Closes #82
Stack