Skip to content

perf(umbp): keep large restore fragments on the gather kernel - #717

Open
isytwu wants to merge 2 commits into
ROCm:mainfrom
isytwu:perf/umbp-gather-large-fragments
Open

isytwu wants to merge 2 commits into
ROCm:mainfrom
isytwu:perf/umbp-gather-large-fragments

Conversation

@isytwu

@isytwu isytwu commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Problem

HbmCopyEngine moves a multi-segment host↔GPU batch either with the gather kernel (one launch per device) or with one blocking hipMemcpy per segment. A batch left the kernel as soon as its mean segment reached 128 KiB (kGatherFragmentThreshold), on the assumption that a copy that large already amortizes its own submission.
With all 8 GPUs copying at once on MI355X, that assumption does not hold host to device. The kernel runs near PCIe line rate (50–55 GB/s per GPU) for fragments from 8 KiB to 64 MiB, while the per-segment fallback takes 6.2× as long at 128 KiB and 1.3–1.6× as long from 1 MiB up. DeepSeek-V4-Pro never reaches the threshold: its restore segments are at most 64 KiB, and an instrumented TP8 serving run at concurrency 128 made no fallback calls. Large per-layer state does reach it, though (Mamba-style SSM state, long-context KV slabs), so the models with the biggest restores were the ones pushed off the fast path.

Change

  • Restores (host → device) stay on the gather kernel whatever the fragment size. Offloads (device → host) still go to hipMemcpy once their mean segment reaches 4 MiB (kGatherD2HFragmentThreshold), the only case where the copy engine measured faster. Single-segment batches are unchanged.
  • The host alias is resolved once per plan. Each segment used to translate its host address with a locked scan of the registered regions, although Plan() has already bounds-checked every segment against the plan's host endpoint. One lookup for the endpoint now yields the alias base for all of them; an endpoint that no single registration covers still resolves segment by segment.
  • The comments in device_gather.h and §10.1 of design-tree-connector-port.md now point at the measured rule instead of the 128 KiB crossover.

Measurements

A standalone HIP benchmark (not part of this PR) runs a copy of GatherFragmentsKernel with the same launch geometry on one 8× MI355X node (2 sockets, GPUs 0–3 on socket 0). Host memory is hugetlb-backed and registered with hipHostRegister. All 8 GPUs copy concurrently; each figure is the time until every GPU has finished, as a median of 10–20 iterations.
Host → device, 256 MiB per GPU, fragments alternating between the two sockets:

fragment gather kernel hipMemcpy per segment (old fallback) hipMemcpyAsync per segment
128 KiB ¹ 1.25 ms 7.78 ms (6.2×) 6.64 ms (5.3×)
1 MiB 4.89 ms 7.93 ms (1.62×) 7.99 ms (1.63×)
4 MiB 4.88 ms 6.35 ms (1.30×) 6.53 ms (1.34×)
16 MiB 4.87 ms 6.52 ms (1.34×) 6.55 ms (1.34×)
64 MiB 4.97 ms 6.56 ms (1.32×) 6.57 ms (1.32×)
¹ 64 MiB per GPU.
A single GPU copying alone sees the two tie at 64 MiB (kernel 4.72 ms, hipMemcpyAsync 4.69 ms). At no measured size is the kernel the slower choice host to device.
Device → host, 256 MiB per GPU, each GPU writing to memory on its own socket:
fragment gather kernel hipMemcpy per segment hipMemcpyAsync per segment
---: ---: ---: ---:
1 MiB 5.41 ms 7.56 ms 7.17 ms
4 MiB 5.46 ms 5.44 ms 5.39 ms
16 MiB 5.58 ms 4.95 ms 4.94 ms
64 MiB 5.59 ms 4.82 ms 4.80 ms
With the writes split across both sockets, the two paths tie from 16 MiB (6.25 vs 6.27 ms), and at 4 MiB the kernel is still 3% ahead (6.19 vs 6.40 ms).
The per-plan alias lookup is a CPU-side saving. In a DeepSeek-V4-Pro TP8 serving run instrumented with UMBP_HBM_COPY_DEBUG, the submit phase, where the lookups happen, went from 0.512 µs per segment in two baseline runs to 0.501 µs. Total restore time did not change measurably.
There is no end-to-end serving run with a large-state model yet; the numbers above cover the copy path only.

Testing

New cases in test_hbm_backend:

  • LargeRestoreFragmentsStayOnTheGatherKernel: host → device batches of 4 × 1 MiB and 3 × 4 MiB each take one kernel launch.
  • LargeOffloadFragmentsGoToTheCopyEngine: device → host, 4 × 1 MiB takes one launch and 3 × 4 MiB takes none.
  • GatherResolvesSegmentsWhenTheEndpointOutgrowsItsRegistration: only half of the host endpoint is registered, so the alias falls back to per-segment lookup. The batch still takes one launch, and the data is checked.
  • GatherKernelRoundTripsScatteredSegments and ConcurrentGatherBatchesStayIndependent (8 threads × 20 rounds): data correctness on the gather path.
    On an 8× MI355X node, all of these pass:
  • test_hbm_backend: 16/16
  • test_transfer_engine: 13/13
  • test_umbp_pool_client_ranges: 27/27, including the existing launch-count checks in GpuRangesUseGatherKernel
  • test_umbp_pool_client_batch_put: 9/9
  • test_peer_pool: 29/29
  • test_page_backend: 36/36
  • test_standalone_shm_ipc: 11 passed, 1 skipped
    pre-commit is clean.

Every segment of a gather batch translated its host address to the device's
alias with a locked scan of the registered regions, although all segments of
a plan lie in the plan's host endpoint (Plan() bounds-checked each one).  One
lookup covering the endpoint now yields the alias base for all of them;
endpoints no single registration covers still resolve segment by segment.

A restore carries thousands of segments, so this removes thousands of lock
round trips per call; the measured effect on a DeepSeek-V4-Pro TP8 restore is
small (submit phase -2% per segment).
A batch whose mean segment reached 128 KiB left the gather kernel for a
blocking hipMemcpy per segment.  With 8 GPUs copying at once on MI355X the
kernel runs near PCIe line rate (50-55 GB/s per GPU) host to device for
fragments from 8 KiB to 64 MiB, and the copy engine takes 1.3-1.6x as long
from 1 MiB up (6.2x at 128 KiB).  Large per-layer state -- Mamba-style SSM
state, long-context KV slabs -- is exactly what produces such fragments, so
restores now stay on the kernel whatever the fragment size.

Device to host the copy engine catches up at about 4 MiB and is up to 14%
faster from 16 MiB, so offload batches whose mean segment reaches 4 MiB still
go to hipMemcpy.  Single-segment batches are unchanged.

Tests cover both rules, the per-plan alias fallback, and a round trip and a
concurrent-submit case on the gather path.
@isytwu isytwu self-assigned this Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant