Skip to content

feat(xen): permit ballooning at 2MiB when possible - #8

Merged
alexandermerritt merged 6 commits into
edera/mainlinefrom
alexm/xen-balloon-2mb
Sep 25, 2026
Merged

alexandermerritt merged 6 commits into
edera/mainlinefrom
alexm/xen-balloon-2mb

Conversation

@alexandermerritt

@alexandermerritt alexandermerritt commented Sep 25, 2026 •

Copy link
Copy Markdown

Closes COR-59

The Xen balloon driver moves memory between a zone and Xen one 4 KiB page at a
time. This series lets it move whole 2 MiB blocks instead, when it can.

  • Growing a zone: ask Xen for whole 2 MiB chunks in one call, instead of
    512 separate pages.
  • Shrinking a zone: give back whole free 2 MiB chunks.
  • PV zones: update page tables in batches, instead of one hypercall per page.
  • Fallback: if a 2 MiB chunk isn't available (memory is fragmented, or the
    zone is close to its limit), use single pages as before. The balloon never
    gets stuck waiting for a block.
  • OOM: when a zone runs out of memory, the balloon can also grow it with
    blocks.

PV guests do not support large pages, so the only benefit we introduce here is using a multicall to batch updates for larger runs.

Turn it off with xen.balloon_superpages=0 on the zone kernel command line.

Results

Same kernel, with and without the feature, shrinking and then growing a zone
by 2 GiB:

|zone | shrink       | grow         |
|-----|--------------|--------------|
| PV  | 1.5x faster  | 3.3x faster  |
| PVH | no change    | 10x faster   |

Deflate (zone grows) is the priority, as it means a guest is waiting on memory, and releasing pages back into the zone is critical for performance. Inflation can happen asynchronously as-needed.

Testing

  • Every commit builds on its own.
  • Booted PV and PVH zones and ran the grow/shrink benchmark.
  • Not yet tested: the out-of-memory path under real memory pressure.

@alexandermerritt

Copy link
Copy Markdown
Author

balloonbench.sh

Attached is the benchmark script I used to evaluate the performance. You'll need to have the kernel rebuilt first.

XENMEM_populate_physmap takes an extent order, but the only wrapper
exposed here pins it to EXTENT_ORDER, the order that covers one native
page and nothing more. Add a variant that takes an additional order, so
a caller can ask Xen to back a run of pages with a single
machine-contiguous allocation rather than one frame per page.

For a PV domain Xen reports back only the base machine frame of each
ordered extent, the rest of the extent following it. For a PVH domain it
reports nothing, as the guest never handles machine frames.

No functional change: xenmem_reservation_increase() becomes a wrapper at
order 0 and the new variant has no callers yet.

Backends: PV, PVH
Direction: deflate (frames move from Xen to the guest)
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
A PV domain's page tables hold machine frames, so moving a page between
the guest and Xen means rewriting its direct map PTE through Xen. The
per-page helpers in drivers/xen/mem-reservation.c make one
HYPERVISOR_update_va_mapping call per page.

Add exported helpers that do the same for a contiguous run in batches,
through HYPERVISOR_mmu_update as xen_remap_pfn() does:

  - xen_zap_contig_pfns() unmaps a run and invalidates its p2m entries,
    for pages about to be handed to Xen.

  - xen_remap_contig_pfns() maps a run onto frames Xen has just
    populated. It writes the p2m with __set_phys_to_machine(), which
    never allocates, so the caller runs xen_prealloc_p2m_range() over
    the run first, before the populate hypercall, while a failure can
    still be backed out of.

Both return the first failure, whether a PTE they could not find or Xen
refused to update, or a p2m entry that could not be written, and leave
what to do about it to the caller. mmu_update reports a failure back to its caller, where a
multicall batch would only log it.

Neither flushes the TLB: the intended caller only maps over PTEs that
were already cleared, and flushes after unmapping itself.

PVH domains never reach these: their page tables hold guest frames,
which do not change when Xen moves the machine frames behind them.

No callers yet.

Backends: PV only
Direction: inflate and deflate
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
A PV domain's page tables hold machine frames, so moving a page between
the guest and Xen means updating its direct map PTE and p2m entry. The
existing helpers for that take parallel arrays of pages and frames, one
entry per page. That fits order-0 extents, where Xen reports a frame per
page, but not an ordered extent, where it reports only the base frame.

Add variants that take a base page, and for mapping a base frame, and
derive the rest of a contiguous run from them, on top of the batched arch
helpers:

  - xenmem_reservation_va_mapping_reset_contig() unmaps a run before its
    frames are handed to Xen.

  - xenmem_reservation_p2m_prealloc() and
    xenmem_reservation_va_mapping_update_contig() map a run in once Xen
    has populated it. The mapping step cannot allocate, so the
    preallocation is a separate call, made before asking Xen to
    populate: it is the only step expected to fail, and failing after
    the populate would strand the frames Xen handed over.

The two directions treat a failed update differently:

  - A failed map is fatal, as in the per-page helper. Callers treat the
    run as mapped once the call returns, and a page with no mapping
    would fault wherever it was next used, far from the cause.

  - A failed unmap only warns, unlike the per-page helper. The frame
    stays mapped, and Xen will not hand it to another domain while it
    is, so the frame leaks but nothing is corrupted.

All three are no-ops outside a PV domain, like the helpers they sit
beside.

No functional change: the new helpers have no callers yet.

Backends: PV only
Direction: inflate and deflate
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
The balloon inflates a page at a time: it takes whatever pages the guest
buddy allocator hands out and returns their frames to Xen one order-0
extent each. Those frames need not form an aligned run, so Xen's
free_heap_pages() cannot merge them back into a 2 MiB buddy.

Add a second unit of trade, a block of 1 << 9 Xen pages: 2 MiB, the
largest extent Xen grants a domain asking on its own behalf. Inflate now
takes a whole free block with one alloc_pages() and parks it on a block
list, so its frames go back to Xen as an aligned run. On PV the block is
unmapped with the batched contiguous-run helper; on PVH there is nothing
to unmap.

The extents submitted stay order-0. XENMEM_decrease_reservation
releases each frame individually whatever the extent order, so an
ordered extent would change nothing there while asking Xen to trust that
the block is machine-contiguous, a claim it cannot check for a PV
domain, where the frame number it is handed is already the machine
frame.

A block is only an option, never a requirement, in either direction:

  - balloon_inflate() takes a block only when the guest has a whole free
    2 MiB run, without reclaiming or compacting to make one, and
    otherwise inflates a page at a time. A guest with plenty of free
    memory can still have no such run, and that must not stop it
    shrinking.

  - Deflate still populates a page at a time and walks only the page
    list, which cannot see blocks. When that list is empty it now breaks
    a parked block onto it, so the guest can always grow back into
    memory it gave up as a block.

The OOM notifier sized its deflate from the page list counters alone,
so a balloon holding only blocks would look empty to it and the OOM
killer would run while the balloon still held memory. It now counts
block pages too, and its deflate reaches them through the block-breaking
path above.

Set xen.balloon_superpages=0 to inflate a page at a time as before.

Backends: PV, PVH
Direction: inflate (frames move from the guest to Xen); deflate only
 changes to consume blocks
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
Deflate still breaks a parked block into pages and asks Xen to populate
each with its own order-0 extent, which lets Xen back the block with any
512 frames it has.

Populate a parked block as a single order-9 extent instead, so that Xen
allocates one machine-contiguous 2 MiB run for it:

  - On PVH Xen maps that run with a single 2 MiB EPT entry rather than
    512 4 KiB ones.

  - On PV there is no EPT, and the guest cannot use 2 MiB mappings. The
    guest maps the run in itself, with the batched contiguous-run
    helper, from the base machine frame Xen reports.

The p2m leaves of each block are preallocated while walking the block
list, before the populate hypercall: mapping the block afterwards cannot
allocate, and once Xen has populated an extent there is no way to hand
it back. Failing there only shortens the batch.

A block is only an option here too. balloon_deflate() falls back to the
page at a time path when a block makes no progress: Xen refuses an
order-9 extent when the host has no free 2 MiB run, or when the domain
is less than 2 MiB below its max_pages ceiling, and neither should stop
the guest growing by what it can still get. Credits smaller than a block
take the page at a time path as well.

Blocks are populated at most BALLOON_BLOCK_BATCH at a time, which bounds
the work done under balloon_mutex; the OOM notifier can only trylock it.

Backends: PV, PVH
Direction: deflate (frames move from Xen to the guest)
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
The OOM notifier deflates through increase_reservation(), which works a
page at a time and reaches parked blocks only by breaking them up.

Deflate through balloon_deflate() instead, so that an OOM deflate
populates a whole block as a single order-9 extent where Xen can supply
one, and still falls back to single pages where it cannot, including the
last 2 MiB below the domain's max_pages ceiling.

Backends: PV, PVH
Direction: deflate (frames move from Xen to the guest)
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
@alexandermerritt

Copy link
Copy Markdown
Author

force pushed with the commits now signed by me

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants