Repository navigation
Conversation
Every segment of a gather batch translated its host address to the device's alias with a locked scan of the registered regions, although all segments of a plan lie in the plan's host endpoint (Plan() bounds-checked each one). One lookup covering the endpoint now yields the alias base for all of them; endpoints no single registration covers still resolve segment by segment. A restore carries thousands of segments, so this removes thousands of lock round trips per call; the measured effect on a DeepSeek-V4-Pro TP8 restore is small (submit phase -2% per segment).
A batch whose mean segment reached 128 KiB left the gather kernel for a blocking hipMemcpy per segment. With 8 GPUs copying at once on MI355X the kernel runs near PCIe line rate (50-55 GB/s per GPU) host to device for fragments from 8 KiB to 64 MiB, and the copy engine takes 1.3-1.6x as long from 1 MiB up (6.2x at 128 KiB). Large per-layer state -- Mamba-style SSM state, long-context KV slabs -- is exactly what produces such fragments, so restores now stay on the kernel whatever the fragment size. Device to host the copy engine catches up at about 4 MiB and is up to 14% faster from 16 MiB, so offload batches whose mean segment reaches 4 MiB still go to hipMemcpy. Single-segment batches are unchanged. Tests cover both rules, the per-plan alias fallback, and a round trip and a concurrent-submit case on the gather path.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
HbmCopyEnginemoves a multi-segment host↔GPU batch either with the gather kernel (one launch per device) or with one blockinghipMemcpyper segment. A batch left the kernel as soon as its mean segment reached 128 KiB (kGatherFragmentThreshold), on the assumption that a copy that large already amortizes its own submission.With all 8 GPUs copying at once on MI355X, that assumption does not hold host to device. The kernel runs near PCIe line rate (50–55 GB/s per GPU) for fragments from 8 KiB to 64 MiB, while the per-segment fallback takes 6.2× as long at 128 KiB and 1.3–1.6× as long from 1 MiB up. DeepSeek-V4-Pro never reaches the threshold: its restore segments are at most 64 KiB, and an instrumented TP8 serving run at concurrency 128 made no fallback calls. Large per-layer state does reach it, though (Mamba-style SSM state, long-context KV slabs), so the models with the biggest restores were the ones pushed off the fast path.
Change
hipMemcpyonce their mean segment reaches 4 MiB (kGatherD2HFragmentThreshold), the only case where the copy engine measured faster. Single-segment batches are unchanged.Plan()has already bounds-checked every segment against the plan's host endpoint. One lookup for the endpoint now yields the alias base for all of them; an endpoint that no single registration covers still resolves segment by segment.device_gather.hand §10.1 ofdesign-tree-connector-port.mdnow point at the measured rule instead of the 128 KiB crossover.Measurements
A standalone HIP benchmark (not part of this PR) runs a copy of
GatherFragmentsKernelwith the same launch geometry on one 8× MI355X node (2 sockets, GPUs 0–3 on socket 0). Host memory is hugetlb-backed and registered withhipHostRegister. All 8 GPUs copy concurrently; each figure is the time until every GPU has finished, as a median of 10–20 iterations.Host → device, 256 MiB per GPU, fragments alternating between the two sockets:
hipMemcpyper segment (old fallback)hipMemcpyAsyncper segmenthipMemcpyAsync4.69 ms). At no measured size is the kernel the slower choice host to device.hipMemcpyper segmenthipMemcpyAsyncper segmentUMBP_HBM_COPY_DEBUG, the submit phase, where the lookups happen, went from 0.512 µs per segment in two baseline runs to 0.501 µs. Total restore time did not change measurably.Testing
New cases in
test_hbm_backend:LargeRestoreFragmentsStayOnTheGatherKernel: host → device batches of 4 × 1 MiB and 3 × 4 MiB each take one kernel launch.LargeOffloadFragmentsGoToTheCopyEngine: device → host, 4 × 1 MiB takes one launch and 3 × 4 MiB takes none.GatherResolvesSegmentsWhenTheEndpointOutgrowsItsRegistration: only half of the host endpoint is registered, so the alias falls back to per-segment lookup. The batch still takes one launch, and the data is checked.GatherKernelRoundTripsScatteredSegmentsandConcurrentGatherBatchesStayIndependent(8 threads × 20 rounds): data correctness on the gather path.On an 8× MI355X node, all of these pass:
test_hbm_backend: 16/16test_transfer_engine: 13/13test_umbp_pool_client_ranges: 27/27, including the existing launch-count checks inGpuRangesUseGatherKerneltest_umbp_pool_client_batch_put: 9/9test_peer_pool: 29/29test_page_backend: 36/36test_standalone_shm_ipc: 11 passed, 1 skippedpre-commit is clean.