Skip to content

feat(umbp): make the heartbeat interval changeable at runtime - #692

Draft
isytwu wants to merge 1 commit into
mainfrom
feat/umbp-runtime-heartbeat-interval
Draft

isytwu wants to merge 1 commit into
mainfrom
feat/umbp-runtime-heartbeat-interval

Conversation

@isytwu

@isytwu isytwu commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

Problem

The heartbeat interval a peer uses is derived on the master from UMBP_HEARTBEAT_TTL_SEC / UMBP_HEARTBEAT_INTERVAL_DIVISOR, and both are read once into function-local statics — they cannot change in a running process. The derived value then reaches a peer only in RegisterClientResponse, so even restarting the master leaves already-connected peers on their old interval until they reconnect.

Retuning reporting cadence therefore meant restarting a master. That is disproportionate for a knob that only affects how often peers report: every peer has to re-register and re-ship a full-sync snapshot before routing is whole again.

This matters because the interval is the backstop for how long a completed BatchPut stays invisible to the master. Events ship either when the peer's outbox hits UMBP_AUTO_FLUSH_EVENT_THRESHOLD (128) or when the heartbeat fires (5s by default). A batch smaller than the threshold waits up to the full interval.

Change

Re-advertise the effective interval on every HeartbeatResponse, and add a SetRuntimeConfig RPC that overrides it in place, driven by a new umbp_admin CLI:

The UMBP_* knobs that seed the heartbeat interval are read once into
function-local statics, so they cannot be changed in a running process.
Worse, the interval only reaches a peer in RegisterClientResponse, so even
restarting the master leaves already-connected peers on their old value
until they reconnect.

Retuning reporting cadence therefore meant restarting a master, which is
disproportionate: every peer must re-register and re-ship a full-sync
snapshot before routing is whole again.

Re-advertise the effective interval on every HeartbeatResponse and add a
SetRuntimeConfig RPC (plus a umbp_admin CLI) that overrides it in place.
A change converges within one old heartbeat period with nothing restarted
and no peer re-registering.

Deliberately narrow: heartbeat_ttl, max_missed_heartbeats and the reaper
are untouched, so failure-detection semantics are unchanged. An override
is clamped to the expiry window, since past that point a peer would be
reaped between two of its own heartbeats -- the invariant the interval
divisor exists to protect. An advertised 0 means "no opinion" (what a
master predating the field sends), and is ignored rather than adopted,
which would collapse the wait and spin the heartbeat thread.

Verified by unit tests (including an inverted run confirming the new case
fails when the client-side adoption is disabled), single-node end-to-end,
and a two-node Slurm job: 5000ms baseline -> 1000ms live -> clamp at
30000ms -> revert to 5000ms, all measured from master-side metrics, with
the peer alive throughout.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@isytwu isytwu self-assigned this Sep 21, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant