Repository navigation
Conversation
The standalone server's client_mu_ is one mutex per NODE -- the server
address is keyed on UMBP_NODE_ID and a bootstrap lock keeps a single
server process, so all 8 TP ranks share it. Two paths held it
exclusively across work that is not short, and every BatchExists on the
node queued behind them:
- RegisterBackendMemory wrapped the inner RegisterMemory, which ends in
an RDMA MR pin over the worker's whole host KV pool. Measured in the
minutes (see RegisterMemoryRpcTimeoutMs's comment: 1187840 MiB in
599.6 s, plus per-GPU IPC registration each well over a minute).
Ranks register at staggered times during warmup, which is what made
the stall a startup-window phenomenon that clears by itself.
- The put handlers held it until the backend call returned, covering
the data copy and the BatchRoutePut round trip to master.
The lock could not simply be narrowed because it doubled as the lifetime
barrier for resolved host/GPU mappings: ReleaseRegisteredMemory took it
exclusively, so nothing could munmap a buffer -- or, more importantly,
tear down its RDMA MR, since PoolClient hands out copies of a region's
TransferRef -- underneath a copy in flight.
Replace the barrier with a per-region pin. memory_ holds shared_ptr to an
immutable region plus an atomic count; ResolveRange/ResolveRanges, the
one choke point every data path goes through, pin what they resolve while
still holding memory_mu_, which is what makes the acquisition race-free.
ReleaseRegisteredMemory waits for the count to drain before deregistering
or unmapping, holding neither lock while it waits. Data handlers and
RegisterBackendMemory now take client_mu_ shared; writes reuse the read
path's ConditionalDataLock so the SSD medium keeps serializing exactly as
before, and only DRAM gains concurrency. Clear and shutdown stay
exclusive.
PoolClient had the same shape one layer down: the slow
transfer_engine_->RegisterMemory ran under registered_mem_mutex_, the
lock FindRegisteredMemory takes once per range on every transfer. Move
that exclusion to a new registration_mutex_ held across the engine call
-- that is what IOEngine actually needs, its memPool and backends
carrying no lock of their own -- and leave registered_mem_mutex_ for the
table insert/erase only.
Pinning until the call returns is sufficient because the inner data-plane
calls are synchronous with respect to caller pointers: the remote
submit-then-wait handles are function-local and always waited before
return, with the handle's destructor draining on exceptional exit.
Tests: a multi-region ranged variant of the deregistration race (its
ranges are listed last-region-first on purpose -- a teardown releases
regions in registration order, so a pin covering only the first region
resolved would otherwise be shielded by that ordering), concurrent puts,
registration churn under load, and a bounded-latency exists probe.
Verified by mutation: removing the pin wait segfaults the pre-existing
deregistration test 3 runs out of 3, and pinning only the first region
segfaults the new ranged test 5 out of 5.
…d thread ConcurrentRegistrationsAndLookupsStayConsistent runs BatchPut on a raw std::thread as a load generator; its sibling tests (StagingFallbackSucceedsWithoutWarn, RegisteredSrcsNoWarn) call the same cross-node BatchPut from the main thread, where gtest's exception handler turns a throw -- e.g. "no active RDMA device on this host" -- into a graceful test failure. An exception escaping a std::thread has no such handler and calls std::terminate, taking down the whole binary instead of just this test. Wrap the call in try/catch, same "not asserted" posture as the plain false result this load generator already ignores.
The write-up was an investigation record, not reference material the repo needs to carry: it is mostly a narrative of how the stall was found, and the parts worth keeping -- why client_mu_ cannot be the lifetime barrier, what the pin replaces it with, and what the regression tests are actually pinning down -- already live in the code comments and the commit that made the change.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
StandaloneServer::client_mu_is one mutex per node, not per rank — the server address is keyed onUMBP_NODE_IDand a bootstrap lock keeps a single server process, so all 8 TP ranks share it. Two paths held it exclusively across work that is not short, and everyBatchExistson the node queued behind them:Memory registration — the dominant cause of the startup stall.
RegisterFd→RegisterBackendMemorywrapped the innerRegisterMemory, which ends in an RDMA MR pin over the worker's whole host KV pool.standalone_process_client.cpprecords the measured cost: "1187840 MiB in 599.6 s", plus per-GPU IPC registration each well over a minute. Ranks register at staggered times during warmup, so for minutes at a stretch the whole node's data plane sat behind one rank's pin. This matches the observed symptom exactly: confined to the startup window, 70-260 s, self-clearing, varying run to run.Bulk writes. The put handlers held it until the backend call returned, covering the data copy and the
BatchRoutePutround trip to master.Why the lock could not simply be narrowed
Holding
client_mu_across the backend call was also, incidentally, the lifetime barrier for resolved host/GPU mappings — nothing couldmunmapa buffer, or tear down its RDMA MR, underneath a copy in flight. (PoolClienthands out copies of a region'sTransferRef, so the inner client has no pin of its own.)The fix
standalone_server.cpp— a per-region pin replaces the lock as the lifetime barrier.memory_holdsshared_ptr<Region>(immutableRegisteredMemory+ atomic count).ResolveRange/ResolveRanges— the single choke point every data path goes through — pin what they resolve while still holdingmemory_mu_, which is what makes the acquisition race-free.ReleaseRegisteredMemorywaits for the count to drain, holding neither lock. Data handlers andRegisterBackendMemorynow takeclient_mu_shared; writes reuse the read path'sConditionalDataLock, so the SSD medium keeps serializing exactly as before and only DRAM gains concurrency.Clearand shutdown stay exclusive.pool_client.cpp— the same shape one layer down: the slowtransfer_engine_->RegisterMemoryran underregistered_mem_mutex_, the lockFindRegisteredMemorytakes once per range on every transfer. Moved that exclusion to a newregistration_mutex_held across the engine call — that is whatIOEngineactually needs, itsmemPool/backendscarrying no lock of their own. Moved, not relaxed.Why the inner client tolerates the new concurrency
DistributedClientalready had exactly this discipline (op_mutex_shared for every data op and register/deregister, exclusive only forClear/Close).PoolClientserializes just its staging arenas;PeerPoolkeeps backend calls outsideoperation_mutex_; each transfer engine carries its own lock. Pinning until the call returns is sufficient because the inner data-plane calls are synchronous w.r.t. caller pointers — the remote submit-then-wait handles are function-local and always waited before return, with the destructor draining on exceptional exit.Testing
test_standalone_shm_ipc15/15, stable over 5 consecutive runs. New cases: multi-region ranged deregistration race, concurrent puts, registration churn under load, bounded-latency exists probe. PlusConcurrentRegistrationsAndLookupsStayConsistentintest_pool_client_batch_put.Verified by mutation — the new tests have teeth:
WaitForPinsZeroreturns immediatelyRegionPins::Addpins only the first regionThe second one is worth calling out: the multi-region test initially passed with that mutation, because a teardown releases regions in registration order and that ordering accidentally shielded the unpinned region. The test now lists its ranges last-region-first on purpose, which removes the accident and leaves the pin as the only thing standing between the copy and an unmapped buffer.
TSAN: 15/15 pass. 67 warnings, all with their peer frame in uninstrumented prebuilt libs (
libgrpc/libprotobuf/libhsa-runtime64); zero frames in any changed function — and those are instrumented, so the negative is meaningful. Running only the pre-existing 11 tests yields 13 of the same class.Known environment dependency:
PoolClientRangesTest.RemoteRoundTrip*,StaleSelfLocationIsExcludedBeforeRemoteFetch,BatchPutWarnTest.StagingFallbackSucceedsWithoutWarnandRegisteredSrcsNoWarnneed a working RDMA device and fail identically with and without this change on a host without one.Relationship to #678
#678 arms these RPCs with a client-side deadline — a bound on an unbounded hang, and explicitly a stopgap ("not a fix for the lock itself"). This PR is that root cause. The two are independent: no file overlap, either can merge first.
The stall was traced to
client_mu_being a per-node lock held exclusively across memory registration and bulk writes; the details are in the commit message ofaef415a4.Validation (independent, at
a874b736— base branch + this fix + the RegisterMemory deadline fix)test_standalone_shm_ipcandtest_pool_client_batch_putboth pass exceptthe 2 tests with a documented RDMA-device environment dependency
(
StagingFallbackSucceedsWithoutWarn,RegisteredSrcsNoWarn), which failidentically with and without this change on a host lacking a working RDMA
device — not a regression.
crsuse2-m2m-v2-015andcrsuse2-m2m-v2-012(Kimi-K3, TP8, DCP8,
CONC=24,ARM=umbp), rebuilding the pinned image atthis commit: the startup-window freeze this fix targets (
aiperf'sreturned/in_flightfrozen for 70–260+ s right afterserver ready,confirmed via
py-spy --nativeto be blocked onclient_mu_) did notreproduce. The unfixed image reliably froze in the same window under the
same recipe; the fixed image kept advancing through it.
surface change; both UMBP call sites in
sglang-k3are unaffected.the first time (this fix removes an incidental throttle the client_mu_ bug
was applying to actual concurrency): an sglang-k3-side mamba-cache
exhaustion (
Can not alloc mamba cache) under sustained long-contextchunked-prefill load. Out of scope for this PR — not caused by it, no file
overlap — tracked separately in
/apps/yutongwu/store/dcp/MAMBA-CACHE-EXHAUSTION-REPORT.md.