Repository navigation
GCS: retry transient upload and token failures - #37
Draft
achudnovskij wants to merge 1 commit into
Draft
achudnovskij wants to merge 1 commit into
achudnovskij wants to merge 1 commit into
Conversation
GCS uploads were a single streamed attempt, so one 503, 429 or connection reset failed the put. S3 buffers bodies under its 32 MiB single-PUT threshold and retries in place; GCS only got RetryingStorage's retries, which need a known size_hint <= 8 MiB. Compressed or encrypted WAL is pushed with no size hint, so in practice no GCS WAL upload was retried. Buffer put and put_if_absent bodies up to 32 MiB (the S3 threshold; every WAL segment fits) and retry transients with the storage retry policy, now threaded through from config like S3's. Bodies over the limit, or that overflow it with no size hint, still stream once: they can't be replayed without buffering the whole thing, and a resumable upload is a larger change. A retried put_if_absent whose earlier attempt landed but lost its response gets 412 and reports AlreadyExists, the same as S3. Also classify 429/5xx from the token endpoint (metadata server or oauth2) as transient Http errors instead of Auth, which is never retried. With node identity the metadata server is hit on every hourly refresh, so a brief blip there failed every operation in that window. The token is fetched inside each upload attempt, so the retry covers it. 4xx stays Auth. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
serprex
approved these changes
Oct 6, 2026
serprex
left a comment
Member
There was a problem hiding this comment.
Thinking we should use backon if we're going to keep adding retry logic, it's what walshadow uses, but that can be followup
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
On GCS, a single transient error fails an upload outright; on S3 the same error is retried. Found while moving wal-listener to GCS (ClickHouse/wal-listener#25). There, a failed DR-tail PUT during failover cuts the tail short, and a failed fallback-archive PUT stalls that pass.
Two gaps, both GCS-only:
GcsStorage::put/put_if_absentstreamed the body once. S3 buffers bodies up to its 32 MiB single-PUT threshold and retries in place. GCS only gotRetryingStorage's retries, which requiresize_hint <= 8 MiB. Compressed or encrypted WAL is pushed withsize_hint = None, so in practice no GCS WAL upload was retried. That covers walruswal-pushand every library user.oauth2.googleapis.combecameStorageError::Auth, whichis_transient()always rejects. The S3 IMDS path instead returnsHttp{status}, so its 5xx is retried. With node identity (GCS: Support Node/VM identity auth. #33) the metadata server is called on every hourly token refresh, so a brief 503 or 429 there failed every GCS operation in that window, reads included.What changes
putandput_if_absentbuffer bodies up toBUFFERED_UPLOAD_LIMIT(32 MiB, the same as S3's threshold) and retry transient errors withwith_retry. With no size hint, it reads one byte past the cap to detect overflow, the same way S3'sputdoes. An overflowing body is chained back together and streamed once, as before. A resumable upload would let those retry too, but it's a larger change and WAL never needs it.GcsStorage::with_retry_policy, asS3Storagedoes).GcsStorage::newkeeps its signature and uses the default policy.StorageError::Http(transient); other 4xx stayAuth. A bad key, missing scope, or missing service account still fails immediately.put_if_absentsemantics are unchanged. If an attempt succeeded but its response was lost, the retry gets 412 and reportsAlreadyExists, the same as S3.Worth a look
RetryingStorageretries around the backend's own retries. S3 already behaves this way, so I kept the two backends consistent rather than fixing it here.Tests
The new tests run against the in-process mock (
test_http::serve):put_retries_transient_failures_for_unknown_size_bodiessize_hint=None→ 3 attempts, body replayed intact; 403 not retried; attempts exhausted → last 503 returnedput_streams_oversized_bodies_in_one_attemptNonehint) and not retried (Somehint)put_if_absent_retries_then_maps_outcomeCreated; 503 then 412 →AlreadyExists;ifGenerationMatch=0sent on every attempttoken_endpoint_errors_are_transient_only_for_throttling_and_5xxAuth; end to end, a metadata server that fails once no longer fails theputMutation check: I reintroduced each bug separately. Forcing the old single streamed attempt fails 3 of the new tests; forcing token errors back to
Authfails the token test.Local gates (rust:1-bookworm):
cargo fmt --check✅,cargo clippy --all-targets --locked -D warnings✅,cargo test --locked✅ (486 lib tests + integration). Not run locally:pg-compatagainst fake-gcs-server; CI runs it.🤖 Generated with Claude Code