Skip to content

feat(xen): permit ballooning at 2MiB when possible - #10

Open
alexandermerritt wants to merge 6 commits into
edera/6.18-ltsfrom
alexm/xen-balloon-2mb-6.18
Open

alexandermerritt wants to merge 6 commits into
edera/6.18-ltsfrom
alexm/xen-balloon-2mb-6.18

Conversation

@alexandermerritt

Copy link
Copy Markdown

Back port of #8

XENMEM_populate_physmap takes an extent order, but the only wrapper
exposed here pins it to EXTENT_ORDER, the order that covers one native
page and nothing more. Add a variant that takes an additional order, so
a caller can ask Xen to back a run of pages with a single
machine-contiguous allocation rather than one frame per page.

For a PV domain Xen reports back only the base machine frame of each
ordered extent, the rest of the extent following it. For a PVH domain it
reports nothing, as the guest never handles machine frames.

No functional change: xenmem_reservation_increase() becomes a wrapper at
order 0 and the new variant has no callers yet.

Backends: PV, PVH
Direction: deflate (frames move from Xen to the guest)
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
(cherry picked from commit 9fe9c09)
A PV domain's page tables hold machine frames, so moving a page between
the guest and Xen means rewriting its direct map PTE through Xen. The
per-page helpers in drivers/xen/mem-reservation.c make one
HYPERVISOR_update_va_mapping call per page.

Add exported helpers that do the same for a contiguous run in batches,
through HYPERVISOR_mmu_update as xen_remap_pfn() does:

  - xen_zap_contig_pfns() unmaps a run and invalidates its p2m entries,
    for pages about to be handed to Xen.

  - xen_remap_contig_pfns() maps a run onto frames Xen has just
    populated. It writes the p2m with __set_phys_to_machine(), which
    never allocates, so the caller runs xen_prealloc_p2m_range() over
    the run first, before the populate hypercall, while a failure can
    still be backed out of.

Both return the first failure, whether a PTE they could not find or Xen
refused to update, or a p2m entry that could not be written, and leave
what to do about it to the caller. mmu_update reports a failure back to its caller, where a
multicall batch would only log it.

Neither flushes the TLB: the intended caller only maps over PTEs that
were already cleared, and flushes after unmapping itself.

PVH domains never reach these: their page tables hold guest frames,
which do not change when Xen moves the machine frames behind them.

No callers yet.

[ 6.18 backport: context conflict only. The helpers go in front of
  xen_exchange_memory(), whose comment still describes the 6.18
  signature taking pfns_in; the added lines are unchanged. ]

Backends: PV only
Direction: inflate and deflate
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
(cherry picked from commit e8f2d6d)
A PV domain's page tables hold machine frames, so moving a page between
the guest and Xen means updating its direct map PTE and p2m entry. The
existing helpers for that take parallel arrays of pages and frames, one
entry per page. That fits order-0 extents, where Xen reports a frame per
page, but not an ordered extent, where it reports only the base frame.

Add variants that take a base page, and for mapping a base frame, and
derive the rest of a contiguous run from them, on top of the batched arch
helpers:

  - xenmem_reservation_va_mapping_reset_contig() unmaps a run before its
    frames are handed to Xen.

  - xenmem_reservation_p2m_prealloc() and
    xenmem_reservation_va_mapping_update_contig() map a run in once Xen
    has populated it. The mapping step cannot allocate, so the
    preallocation is a separate call, made before asking Xen to
    populate: it is the only step expected to fail, and failing after
    the populate would strand the frames Xen handed over.

The two directions treat a failed update differently:

  - A failed map is fatal, as in the per-page helper. Callers treat the
    run as mapped once the call returns, and a page with no mapping
    would fault wherever it was next used, far from the cause.

  - A failed unmap only warns, unlike the per-page helper. The frame
    stays mapped, and Xen will not hand it to another domain while it
    is, so the frame leaks but nothing is corrupted.

All three are no-ops outside a PV domain, like the helpers they sit
beside.

No functional change: the new helpers have no callers yet.

Backends: PV only
Direction: inflate and deflate
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
(cherry picked from commit cab7ddd)
The balloon inflates a page at a time: it takes whatever pages the guest
buddy allocator hands out and returns their frames to Xen one order-0
extent each. Those frames need not form an aligned run, so Xen's
free_heap_pages() cannot merge them back into a 2 MiB buddy.

Add a second unit of trade, a block of 1 << 9 Xen pages: 2 MiB, the
largest extent Xen grants a domain asking on its own behalf. Inflate now
takes a whole free block with one alloc_pages() and parks it on a block
list, so its frames go back to Xen as an aligned run. On PV the block is
unmapped with the batched contiguous-run helper; on PVH there is nothing
to unmap.

The extents submitted stay order-0. XENMEM_decrease_reservation
releases each frame individually whatever the extent order, so an
ordered extent would change nothing there while asking Xen to trust that
the block is machine-contiguous, a claim it cannot check for a PV
domain, where the frame number it is handed is already the machine
frame.

A block is only an option, never a requirement, in either direction:

  - balloon_inflate() takes a block only when the guest has a whole free
    2 MiB run, without reclaiming or compacting to make one, and
    otherwise inflates a page at a time. A guest with plenty of free
    memory can still have no such run, and that must not stop it
    shrinking.

  - Deflate still populates a page at a time and walks only the page
    list, which cannot see blocks. When that list is empty it now breaks
    a parked block onto it, so the guest can always grow back into
    memory it gave up as a block.

The OOM notifier sized its deflate from the page list counters alone,
so a balloon holding only blocks would look empty to it and the OOM
killer would run while the balloon still held memory. It now counts
block pages too, and its deflate reaches them through the block-breaking
path above.

Set xen.balloon_superpages=0 to inflate a page at a time as before.

Backends: PV, PVH
Direction: inflate (frames move from the guest to Xen); deflate only
 changes to consume blocks
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
(cherry picked from commit 534e37a)
Deflate still breaks a parked block into pages and asks Xen to populate
each with its own order-0 extent, which lets Xen back the block with any
512 frames it has.

Populate a parked block as a single order-9 extent instead, so that Xen
allocates one machine-contiguous 2 MiB run for it:

  - On PVH Xen maps that run with a single 2 MiB EPT entry rather than
    512 4 KiB ones.

  - On PV there is no EPT, and the guest cannot use 2 MiB mappings. The
    guest maps the run in itself, with the batched contiguous-run
    helper, from the base machine frame Xen reports.

The p2m leaves of each block are preallocated while walking the block
list, before the populate hypercall: mapping the block afterwards cannot
allocate, and once Xen has populated an extent there is no way to hand
it back. Failing there only shortens the batch.

A block is only an option here too. balloon_deflate() falls back to the
page at a time path when a block makes no progress: Xen refuses an
order-9 extent when the host has no free 2 MiB run, or when the domain
is less than 2 MiB below its max_pages ceiling, and neither should stop
the guest growing by what it can still get. Credits smaller than a block
take the page at a time path as well.

Blocks are populated at most BALLOON_BLOCK_BATCH at a time, which bounds
the work done under balloon_mutex; the OOM notifier can only trylock it.

Backends: PV, PVH
Direction: deflate (frames move from Xen to the guest)
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
(cherry picked from commit 0fe1389)
The OOM notifier deflates through increase_reservation(), which works a
page at a time and reaches parked blocks only by breaking them up.

Deflate through balloon_deflate() instead, so that an OOM deflate
populates a whole block as a single order-9 extent where Xen can supply
one, and still falls back to single pages where it cannot, including the
last 2 MiB below the domain's max_pages ceiling.

Backends: PV, PVH
Direction: deflate (frames move from Xen to the guest)
Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
(cherry picked from commit ee71370)
@alexandermerritt
alexandermerritt force-pushed the alexm/xen-balloon-2mb-6.18 branch from 3907770 to b553de6 Compare September 28, 2026 21:32
@alexandermerritt

Copy link
Copy Markdown
Author

Here's the range-diff between the two:

1:  9fe9c09f2a784 ! 1:  5373577968b26 xen/mem-reservation: add ordered populate
    @@ Commit message
         Backends: PV, PVH
         Direction: deflate (frames move from Xen to the guest)
         Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
    +    (cherry picked from commit 9fe9c09f2a78466ee6a38b580537b7cc3cf45bb7)

      ## drivers/xen/mem-reservation.c ##
     @@ drivers/xen/mem-reservation.c: void __xenmem_reservation_va_mapping_reset(unsigned long count,
2:  e8f2d6dcfc604 ! 2:  f259909724f86 x86/xen: add batched helpers for remapping a contiguous pfn run
    @@ Commit message

         No callers yet.

    +    [ 6.18 backport: context conflict only. The helpers go in front of
    +      xen_exchange_memory(), whose comment still describes the 6.18
    +      signature taking pfns_in; the added lines are unchanged. ]
    +
         Backends: PV only
         Direction: inflate and deflate
         Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
    +    (cherry picked from commit e8f2d6dcfc6042881d74c1bd6d1daa8327221bf8)

      ## arch/x86/include/asm/xen/page.h ##
     @@ arch/x86/include/asm/xen/page.h: extern unsigned long  xen_max_p2m_pfn;
    @@ arch/x86/include/asm/xen/page.h: extern unsigned long  xen_max_p2m_pfn;

      ## arch/x86/xen/mmu_pv.c ##
     @@ arch/x86/xen/mmu_pv.c: static void xen_remap_exchanged_ptes(unsigned long vaddr, int order,
    -   xen_mc_issue(true);
    +   xen_mc_issue(0);
      }

     +/*
    @@ arch/x86/xen/mmu_pv.c: static void xen_remap_exchanged_ptes(unsigned long vaddr,
     +EXPORT_SYMBOL_GPL(xen_zap_contig_pfns);
     +
      /*
    -  * Perform the hypercall to exchange a region of our pages to point to memory
    -  * with the required contiguous alignment.  Takes as input the mfns to trade
    +  * Perform the hypercall to exchange a region of our pfns to point to
    +  * memory with the required contiguous alignment.  Takes the pfns as
3:  cab7ddd3e7ee6 ! 3:  fe3dfec6697f9 xen/mem-reservation: add contiguous mapping helpers
    @@ Commit message
         Backends: PV only
         Direction: inflate and deflate
         Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
    +    (cherry picked from commit cab7ddd3e7ee67c9850bb9ce8aa9d33d6a98ead3)

      ## drivers/xen/mem-reservation.c ##
     @@ drivers/xen/mem-reservation.c: void __xenmem_reservation_va_mapping_reset(unsigned long count,
4:  534e37ae697f1 ! 4:  d0fb3f957141e xen/balloon: inflate in 2 MiB blocks
    @@ Commit message
         Direction: inflate (frames move from the guest to Xen); deflate only
          changes to consume blocks
         Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
    +    (cherry picked from commit 534e37ae697f1b41bcf4087dfe75e5fb1b9fe022)

      ## drivers/xen/balloon.c ##
     @@ drivers/xen/balloon.c: static const struct ctl_table balloon_table[] = {
5:  0fe138949be10 ! 5:  26f803bc63df7 xen/balloon: deflate in 2 MiB blocks
    @@ Commit message
         Backends: PV, PVH
         Direction: deflate (frames move from Xen to the guest)
         Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
    +    (cherry picked from commit 0fe138949be10b99c4bcbbe26672267725d1254a)

      ## drivers/xen/balloon.c ##
     @@ drivers/xen/balloon.c: static const struct ctl_table balloon_table[] = {
6:  ee71370a693c1 ! 6:  b553de63659a8 xen/balloon: deflate 2 MiB blocks in the OOM path
    @@ Commit message
         Backends: PV, PVH
         Direction: deflate (frames move from Xen to the guest)
         Signed-off-by: Alexander M. Merritt <alexander@edera.dev>
    +    (cherry picked from commit ee71370a693c12056b7c44cdd4f7b2709d227e35)

      ## drivers/xen/balloon.c ##
     @@ drivers/xen/balloon.c: static int balloon_oom_notify(struct notifier_block *nb, unsigned long dummy,

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant