Build without std, measure the lock choice, and 0.0.13 - #18
Merged
Merged
Conversation
added 9 commits
September 7, 2026 17:42
`concurrent` is the only reason this crate needs `std`, its 31 lock sites are
parking_lot, and `cdc` and `multimap` both imply `concurrent`. So a consumer
wanting only `ChangeEvent` and `Pair` takes parking_lot and libc with them.
DataBucket is exactly that consumer, and it is what blocks its format layer
from being no_std.
Swapping to a spinlock is the obvious answer and the objection is equally well
known: a spinlock burns a core rather than sleeping, which is why parking_lot
exists. Both are true, in different regimes, and nothing here measured which
one this crate is in. parking_lot does not measure it either: its own tests are
two RwLock regression files and assert nothing about parking or CPU.
Measured on 16 cores, medians of 5, with CPU time from getrusage because wall
clock cannot see a burnt core:
this crate's shape, RwLock<BTreeMap<T, Arc<Mutex<Node>>>>, 200k lookups
threads parking_lot cpu spin cpu spin vs now
16 783.1 ms 391.3 ms -50%
64 805.6 ms 344.2 ms -57%
one lock, held 20000 iterations, oversubscribed
threads parking_lot cpu spin cpu spin vs now
16 16.3 ms 116.0 ms +612%
32 35.3 ms 291.9 ms +727%
Short critical sections favour spinning, because parking costs a syscall pair
per contention event and this crate's node access is a few operations on a key
array. A long hold under oversubscription does not, because a spinner holds the
core the lock holder needs to finish, and the wait extends the thing it waits
on.
**Every measurement is one body, generic over `R: RawMutex`**, instantiated
with parking_lot's raw lock and spin's. That is the proposed shape rather than
a benchmarking convenience: both are lock_api underneath, `ArcMutexGuard<R, T>`
is lock_api's type in both cases, and `lock_arc` comes from lock_api rather
than from parking_lot. So `concurrent` can be generic over `R` and let the
consumer choose, instead of naming a lock and deciding for everyone.
The null column is parking_lot against itself and is the floor for believing
any of it. It sits at 1.00x everywhere except the 128-thread contended row,
where it reads 1.60x: that row is not evidence and is left in rather than
quietly dropped.
…claimed
The first version of this bench varied the index RwLock and the node Mutex at
the same time, then reported the difference as a node-lock result. It is not:
with the index lock held constant every node lock lands inside the null.
this crate's shape, 16 threads, CPU ms
parking_lot 774.8 spin 782.6 spin+yield 755.5 bounded 807.7 null 1.00x
So the node lock does not matter here and the earlier numbers quoted from this
file, -56% and -50% and 372 against 819, were measuring the index lock. The
index RwLock is the bottleneck and is what deserves the next look.
Where a lock does decide something is a long hold under oversubscription:
one lock held 20000 iterations, CPU ms
threads parking_lot spin spin+yield bounded futex
128 153.2 5272.1 987.9 986.9 1680.5
Two things came out of reading parking_lot rather than guessing at it. Its
SpinWait is three exponentially growing pauses, then seven yields, then park,
not the flat 64 spins this file used; copying that schedule made the bounded
arm and the yield arm identical, 986.9 against 987.9, which says the schedule
was never the difference. And the residue is structural: when the schedule runs
out there is nothing to park on without an operating system, so a no_std waiter
keeps running. 6.4x is what that costs.
The futex arm is wrong and is left in only so the next person does not repeat
it. 1680.5 loses to a spinning lock, which a blocking lock cannot legitimately
do. Two rewrites have not fixed it.
Both tables now run the same five locks so a column can be read across them,
and every arm goes through lock_api, which the parking_lot pair shows is free.
WFE idles a core until its event register is set; SEV sets it. ARM documents this as the intended spinlock construction, and Rust never emits it: core::hint::spin_loop() on aarch64 is __isb(SY), confirmed by disassembly. There is no WFE in core::hint and no wrapper in core::arch::aarch64, which is structural rather than an oversight, because WFE is only useful if the releaser pairs it with SEV and a one-sided hint cannot express a protocol. Measured per instruction on this machine: nop 0.3 ns, isb 8.6 ns, wfe 1336.7 ns. So WFE is not a userspace no-op here, and it does not block forever either: it idles about 1.3 us and returns unprompted. **And it does not help, which is the point of adding it.** WFE idles the core but does not yield to the operating system, so a waiter is still a scheduled thread and the lock holder still cannot get a core. At 128 threads on 16 cores this arm costs 5584 ms of CPU against 1241 for a yielding spinner. Below the core count it is level with plain spinning and no better. An earlier version restarted its spin schedule after each WFE, so it yielded between idles, and that accident is what briefly made it look competitive. The rule that falls out: WFE is for threads at most cores, with no operating system to yield to. That is bare metal and pinned threads, not an oversubscribed host, and it is written into the arm's doc comment along with the numbers. Kept rather than deleted because EKOPathRS is a compiler, and a compiler that recognises a spin loop can emit WFE where LLVM does not. This measures what that would be worth and, more usefully, when it would be wrong.
It is not in NAMES and nothing calls it, so clippy was right to flag it. Kept behind a named module rather than deleted, because the useful thing about it is that it is wrong: it loses to a spinning lock, which a blocking lock cannot honestly do, and two rewrites have not fixed it.
The benchmark prices five locks. Nothing checked that the machine underneath it still behaves the way the numbers assume, so a toolchain or hardware change would have quietly turned it into a measurement of something else. Five tests, two in the fast tier. That WFE and SEV execute at all in userspace, and that a waiter is never stuck when no SEV arrives, since both are what make a WFE lock safe to write. Behind --ignored: that WFE actually idles rather than being a nop, that core::hint::spin_loop is not WFE, and the rule itself, that WFE does not reach the scheduler. The header carries what was learned rather than only what is asserted: the instruction and where it exists, that Rust emits __isb(SY) and never WFE and why that is structural, the per-instruction costs, the rule that falls out, and the two findings from the same work that should not be re-derived, that lock_api is free and that the node lock is not this crate's bottleneck. It also records that the negative result is macOS-specific. Under KVM, WFE traps and the hypervisor yields the vCPU, which supplies the half that is missing here, so wfe_does_not_yield_to_the_scheduler is the test to watch on Graviton. If it starts passing comfortably there, the guidance in this file is wrong for that host and needs revisiting.
A B-tree over `alloc` collections and an `ftree` that is already no_std. The
only thing here that ever needed an operating system was one `yield_now`.
::core:: ops, fmt, slice, hint, borrow, marker, mem, iter, hash, cmp, sync::atomic
alloc:: Arc, Vec, String, BTreeMap, vec!, format!
The leading `::` is not decoration. This crate has its own module named
`core`, so inside it `use core::x` resolves to `crate::core::x`.
Four things wanted more than a rename:
- `yield_now` in the two root-publication spin loops falls back to
`spin_loop` without `std`, which is what the branch above it already does.
- `published_keys` was a `HashMap<usize, T>`, keyed by a slot number and
never iterated for order. It is a `BTreeMap` now, the same map without a
hasher, and the one `alloc` has.
- `ftree` is no_std, but its `serde` feature takes serde with default
features, which is `std`. Asking for it unconditionally made every build a
std build. It now arrives with this crate's `serde` feature, which says
`std` out loud, as does `superslice-binary-search`: superslice is not a
no_std crate.
- `parking_lot` is now `parking_lot_lite_hack`, kept under the
`parking_lot` name because that is what the code says. This crate uses
`Mutex`, `RwLock` and `RawMutex`, and the fork has all three. The bench's
own copy of upstream parking_lot is renamed `parking_lot_upstream`,
because one name cannot mean both and that arm is deliberately the real
one, with its std futex.
Doc examples still say `std::`: a doctest compiles as a consumer.
cargo test --all-features 127 + 1 + 2 + 107 doctests passed, 0 failed
clippy, std and no_std clean
x86_64-unknown-none the core builds; `concurrent` stops at
ps-reclaim, whose TLS needs an OS
`"^0.1, >=0.1.4"` and `"0.1.4"` describe the same set, and `crossbeam-skiplist` carried a `"^0.1"` that a bare `"0.1"` already means. One form, used once.
0.12.5 depended on its own core without `default-features = false`, so the
core linked `std` in every build and this crate's no_std column was false one
level down. 0.12.6 is that fix, published today.
--no-default-features parking_lot_lite_hack [arc_lock,send_guard] core []
--features std parking_lot_lite_hack [arc_lock,send_guard,std] core [std]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two things, both of which were sitting unpushed locally. Seven commits, linear.
1. no_std
A B-tree over
alloccollections and anftreethat is alreadyno_std. The only thing here that ever needed an operating system was oneyield_now.The leading
::is not decoration. This crate has its own module namedcore, so inside ituse core::xresolves tocrate::core::x.Four things wanted more than a rename:
yield_nowin the two root-publication spin loops falls back tospin_loop()withoutstd— the same instruction the branch directly above it already executes.published_keyswas aHashMap<usize, T>, keyed by a slot number and never iterated for order. It is aBTreeMapnow: the same map without a hasher, and the oneallochas.ftreeisno_std, but itsserdefeature takes serde with default features, which isstd. Asking for it unconditionally made every build a std build. It now arrives with this crate's ownserdefeature, which saysstdout loud — as doessuperslice-binary-search, because superslice is not ano_stdcrate.parking_lotis nowparking_lot_lite_hack, kept under theparking_lotname because that is what the code says. This crate usesMutex,RwLockandRawMutex, and the fork has all three.Doc examples still say
std::, deliberately: a doctest compiles as a consumer, and a consumer has one.What no_std does not mean here
Not bare metal. The core builds for
x86_64-unknown-none;concurrentstops atps-reclaim, whose participant slot lives in thread-local storage and whoseno_stdpath is pthread keys.no_stdis a statement about the standard library, not about a class of hardware.2. The lock-choice benchmark
benches/locks.rs, five commits, already written and never pushed. It prices this crate's lock choice againstlock_api+parking_lot,spin, and a real futex, in both the uncontended and the held regimes, with a control so the numbers can be checked. One arm is quarantined rather than deleted, with a note saying why it is broken, and a WFE arm records where that instruction is and is not the right one.The bench's own copy of upstream parking_lot is renamed
parking_lot_upstream: the library'sparking_lotis now the fork, one name cannot mean both, and that arm is deliberately the real one with its std futex.An untracked
benches/three.rswas left out. It is a superseded draft of the same measurement and does not compile — it importsarc_swap, which is not a dependency of this crate.Checks
Version 0.0.13.
Companion to pathscale/congee-wt#6 (same port) and pathscale/parking_lot_lite_hack#3.
worktable-vecis what wanted both: pathscale/WorkTable-vec#4 currently declares that itscongeeandwtifeatures implystd, and these two PRs are what removes that.