Skip to content

Spend the maintenance interval between passes, not before the first - #49

Merged
tobi merged 1 commit into
tobi:mainfrom
zaeku:maintain-interval-between-passes
Sep 11, 2026
Merged

tobi merged 1 commit into
tobi:mainfrom
zaeku:maintain-interval-between-passes

Conversation

@zaeku

@zaeku zaeku commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Moves the maintenance.interval sleep from the beginning to the end of the loop body.

maintenance.interval currently sleeps before the first pass, causing newly started processes to sit idle for one interval. Ephemeral maintainers (e.g., serverless hosts or per-request containers) often terminate before executing any work without raising an error.

Moving sleep to the end of the loop fixes this while keeping the drain check at the top.

Trade-off: Simultaneous restarts will no longer be staggered across the interval. Since leases/ guards heavy operations, this results in lock contention rather than duplicate work. I can put this behind a feature flag if default staggering is preferred.

The field is documented as the pause between passes; the loop slept it at the
top of the body, so a freshly started maintainer idled for one interval and a
maintainer that did not live that long never ran a pass at all. The draining
check stays at the top of the loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tobi
tobi merged commit 6465bf5 into tobi:main Sep 11, 2026
@0bserver07

Copy link
Copy Markdown
Contributor

Ran this here: cargo test -p walgit-server --test maintain --test drain passes, and the drain check staying at the top keeps D31 intact.

Two things worth stating in the PR, since they're the actual trade-off:

  • The first pass now starts before prewarm has warmed anything. run_pass opens with registry.list() and then one unit per assigned repo, and on the host that maintains a large repository that unit can be a full Serve sync or a bundle build, so it now races startup instead of getting one interval of serving-only warmth. Might be fine, might be what the leading sleep was for; the file's history doesn't say (it's in the initial release).
  • On simultaneous restarts, compaction and bundle builds are behind leases, but the checkpoint unit isn't (maintain.rs:588), so two hosts would both write the same checkpoint. Idempotent by D22, so wasted round trips rather than duplicate work, but this repo counts round trips.

follow::run_loop has the same sleep-first shape (follow.rs:39), so whichever way this lands, that loop probably wants the same treatment.

@zaeku
zaeku deleted the maintain-interval-between-passes branch September 11, 2026 05:03
brightsparc added a commit to introspection-org/walgit that referenced this pull request Sep 12, 2026
Upstream main (tobi#42-tobi#45, tobi#49) plus the Azure and IRSA branches as
updated for PRs tobi#50 and tobi#48, plus the store-plugin and notify-transport
lines the production image was already built from (a6ce265). The two
conflicts were additive: the Event Grid handshake test and its doc note
now sit beside the loopback notify-transport test and its bullet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T6UZ2AaFQ6m7xgdZKRyHEV
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants