Skip to content

RSI: verify durable resource-aware scheduling across restarts #5

Description

@w4ffl35

Implemented

PR #14 added a PostgreSQL queue with resource-class and concurrency-key admission caps, persisted job states, cancellation of queued work, retry limits, worker heartbeats, and fenced leases. Disposable local PostgreSQL/RustFS tests used 12 simulated workers and 100 jobs; a same-gateway live run observed two evaluation leases overlapping.

Remaining work

Test queued, leased, completed, failed, and cancelled jobs after process restart. Preserve diagnostics from failed attempts across retries and verify admission and recovery across independent concurrent RSI runs. The live observation does not qualify cross-host or cross-gateway operation. Campaign lineage and total budgets are tracked separately in #19.

Acceptance criteria

  • Resource limits prevent oversubscription across concurrent RSI runs.
  • Queued, running, completed, failed, and cancelled jobs survive restart.
  • Resume does not duplicate completed work or lose failed-job diagnostics.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions