Skip to content

E2E flake: post-restore route HTTP 503 triggers missing proxy pod failure #2472

Description

@coderabbitai

Summary

Track a reported post-restore route timing flake in MySQL application two Vol CSI. A transient HTTP 503 causes endpoint resolution to fall back immediately to a proxy pod. When no pod with label curl-tool=true exists, verification fails before the existing healthz retry loop runs.

Requested by @kaovilai. This is a test-infrastructure follow-up to PR #2443, not a request to change its CA handling.

Reported occurrence

  • Job: pull-ci-openshift-oadp-operator-oadp-dev-5.0-e2e-test-aws
  • Build ID: 2105330103645900800
  • Base branch: oadp-dev
  • Cluster: OpenShift 5.0.0-0.nightly-2026-09-29-171841, AWS, linux/amd64
  • Test: [It] Backup and restore tests > Backup and restore applications > MySQL application two Vol CSI
  • Results: 26 passed, 1 failed, 6 pending, 32 skipped.
  • Duration: 233 seconds, with 2 retries.

The supplied analysis reports that backup and restore completed successfully. At 18:02:12 UTC, the restored deployment was Available and its pod was Ready, but the route returned HTTP 503. Proxy fallback then failed with:

Error getting pod for the proxy command: no Pod found

The suspected cause is asynchronous OpenShift router endpoint synchronization after pod readiness. The job logs were not independently inspected when this issue was created. The reporter identifies this as a flake unrelated to PR #2443 and reports that the CA tests passed.

Confirmed code behavior

In the inspected PR branch:

  • tests/e2e/lib/apps.go: getAppEndpointURLAndProxyParams() calls GetRouteEndpointURL() once, then immediately searches for a curl-tool=true proxy pod if that call fails.
  • VerifyBackupRestoreData() returns endpoint-resolution errors before reaching the post-restore healthz loop.
  • The healthz loop retries up to five times, but cannot handle failures that occur during endpoint resolution.
  • tests/e2e/backup_restore_suite_test.go defines the affected two-volume MySQL CSI test.
  • tests/e2e/lib/flakes.go contains the OADP-5086 deployment-readiness pattern, but no route HTTP 503 pattern was found. This report differs from deployment readiness failure.

Required changes

  1. Add bounded retries for transient route unavailability before proxy fallback, or separate endpoint resolution from reachability probing so the existing verification retries can handle transient HTTP 503 responses.
  2. Preserve proxy fallback for environments where direct route access is unavailable.
  3. Preserve a useful route error when proxy fallback also fails. Log retry attempts and the final failure.
  4. Add regression coverage for transient HTTP 503 followed by success, persistent route failure, and unavailable proxy pods.
  5. Consider a narrowly scoped flake classification in tests/e2e/lib/flakes.go linked to this issue. Do not classify all route failures or HTTP 503 responses as harmless.

Acceptance criteria

  • A route that returns transient HTTP 503 responses and then becomes reachable within the retry window does not fail solely because no proxy pod exists.
  • Persistent route failures still fail within a bounded interval with useful diagnostics.
  • Existing proxy fallback remains functional.
  • Regression tests cover success after retry and exhausted retries.
  • Repeated runs of the affected MySQL two-volume CSI scenario validate the fix.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions