Skip to content

Epoch-based reclamation and per-page write locking - #77

Merged
pathscale merged 10 commits into
masterfrom
audit/epoch
Sep 1, 2026
Merged

pathscale merged 10 commits into
masterfrom
audit/epoch

Conversation

@pathscale

Copy link
Copy Markdown
Owner

The two architectural items from the 2026-08-31 audit's performance tier, built and benchmarked. Retires four documented known issues: the SeqCst global read counter, the zero-reader reclamation requirement (with its inline backlog-drain spikes), the table-global write lock, and the partition retire-list leak through shared routers.

Phase A: epoch-based reclamation

Readers pin a per-table epoch domain (crossbeam-epoch, already transitive via indexset; per-thread handle cache so a pin is thread-local work) instead of a shared SeqCst RMW. Retired links, pages, and publications defer through a FIFO whose expired prefix runs the existing capacity-recycle logic — reclamation recycles into empty_links/empty_pages, it does not merely free — draining incrementally (≤256/call), never as a stop-the-world backlog. The partition module's remove() defers through the same scheme, and a new collect(&self) reclaims through the shared Arc router, fixing the documented permanent leak; gc(&mut self) remains as the exhaustive variant.

Phase B: per-page write locking

Each data page carries its own barrier; the table-global page_access is deleted. The survey found exactly one genuinely multi-page mutation — the vacuum row move, which takes both barriers in ascending page-id order — and the full lock order is documented on DataPages and shown acyclic. One additional real race closed en route: a maximally stale appender could write through a popped link into a page vacuum had reclaimed and reused; the append path now re-checks current_page_id under the page barrier.

Evidence

  • The reader-overlap reclamation test was verified FAILING on the old scheme before the rewrite; churn probes hold page counts bounded (15-17 pages) where the old scheme grew 66-100.
  • Miri (Tree Borrows) clean over the partition suite, epoch module, and targeted pages tests — and it caught a real bug in the first handle-cache design during development (fixed; the discipline works).
  • Benchmarks on this exact base, M4 Max: reads 9.5 / 14.2 / 15.1 / 15.2-16.4M ops/s at 1/2/4/8 threads (old counter: 9.4 / 8.3 / 8.4 / 7.5M — negative scaling); disjoint-page in-place writes 5.8M at 8 threads vs 3.4M under the global lock; loom partition models 6/6 unchanged.
  • Full workspace matrix green on default / versioned-row-publication / all-features; clippy -D warnings clean on both configs. Semantic overlap with insert_many's batch ghost-staging reviewed: it mutates only through the per-page-barriered DataPages API and its stripes sit at the documented lock-order position.

Semver note (the reason this PR carries NO version bump)

ReadGuard is now !Send: a lazy select iterator held across .await in a spawned task fails to compile instead of silently stalling reclamation. Point reads are unaffected. Documented on the type and in docs/versioned-row-publication.md. Since a master version bump now auto-publishes, the bump to the next version is a separate, deliberate commit once this has soaked.

meh added 10 commits September 1, 2026 08:03
…tion

Readers pin the table's epoch domain instead of a SeqCst RMW on one shared
active_readers cache line. Retired links, pages, and publications enter one
FIFO queue; each retirement defers an epoch marker, and the count of executed
markers releases a safe prefix of the queue for recycling. Reclamation runs
the same capacity-recycle logic as before (empty_links, empty_pages,
publication-shard removal), drains bounded batches instead of the whole
backlog inline, and no longer needs a global zero-reader instant, so deletes
keep reclaiming space under continuously overlapping reads.
The DataPages-level test hands one retired link across two reader threads so
no zero-reader instant ever occurs; it fails against the old counter scheme
and passes with epoch grace. The leak-probe variant churns 5000 updates while
three reader threads scan without pause and asserts page count stays bounded.
remove() defers each removal's drop behind an epoch marker in the set's own
domain; slot readers pin around raw-pointer access, partition_ref returns a
pinned borrow (PartRef), and collect() frees the expired prefix through
&self. remove and get_or_create collect opportunistically, so an Arc-shared
router no longer leaks every removed partition; gc(&mut self) stays as the
exhaustive variant. Loom builds keep the pre-epoch retire-list semantics:
crossbeam-epoch cannot run under loom, and the slot publication protocol the
models check is unchanged in both builds.
…p last

If a cached LocalHandle holds the last Collector reference, crossbeam's
Local::finalize drops the Global from inside the Local's own method frame and
deallocates the Local while a protected reference to it is on the stack;
Miri (Tree Borrows) flags the deallocation as undefined behavior. Caching an
explicit Collector clone that outlives the handle tears the global down from
outside any Local frame instead.
Each Data page now carries its own reader/writer lock; row writes, in-place
updates, hydration, CDC byte capture, and reset-on-reuse serialize only
against access to the same page. The vacuum row move, the one genuinely
multi-page mutation, takes both pages' barriers in ascending page-id order.
The insert append path re-checks current_page_id under the page barrier: a
page never becomes current again once switched away, so the re-check proves
the locked page cannot be in or headed for the empty-page pool, closing the
window where a maximally stale appender could write into a page that was
vacuumed and reused after the unlocked load. Lock order against the row
locks, empty_pages, the pages vector, publication shards, and the empty-link
registry is documented on DataPages and stays acyclic.
@pathscale
pathscale merged commit 8bebac2 into master Sep 1, 2026
11 of 12 checks passed
@pathscale
pathscale deleted the audit/epoch branch September 1, 2026 04:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant