Skip to content

Build without std, measure the lock choice, and 0.0.13 - #18

Merged
pathscale merged 9 commits into
masterfrom
feat/no-std
Sep 7, 2026
Merged

pathscale merged 9 commits into
masterfrom
feat/no-std

Conversation

@pathscale

Copy link
Copy Markdown
Owner

Two things, both of which were sitting unpushed locally. Seven commits, linear.

1. no_std

A B-tree over alloc collections and an ftree that is already no_std. The only thing here that ever needed an operating system was one yield_now.

::core::  ops, fmt, slice, hint, borrow, marker, mem, iter, hash, cmp, sync::atomic
alloc::   Arc, Vec, String, BTreeMap, vec!, format!

The leading :: is not decoration. This crate has its own module named core, so inside it use core::x resolves to crate::core::x.

Four things wanted more than a rename:

  • yield_now in the two root-publication spin loops falls back to spin_loop() without std — the same instruction the branch directly above it already executes.
  • published_keys was a HashMap<usize, T>, keyed by a slot number and never iterated for order. It is a BTreeMap now: the same map without a hasher, and the one alloc has.
  • ftree is no_std, but its serde feature takes serde with default features, which is std. Asking for it unconditionally made every build a std build. It now arrives with this crate's own serde feature, which says std out loud — as does superslice-binary-search, because superslice is not a no_std crate.
  • parking_lot is now parking_lot_lite_hack, kept under the parking_lot name because that is what the code says. This crate uses Mutex, RwLock and RawMutex, and the fork has all three.
default                  WorkTablesIndex [default,std]  parking_lot_lite_hack [arc_lock,send_guard,std]  ps-reclaim [libc,spin,std]
--no-default-features    WorkTablesIndex [concurrent]   parking_lot_lite_hack [arc_lock,send_guard]     ps-reclaim [libc,spin]

Blocked on pathscale/parking_lot_lite_hack#3 being published. That fork depends on its own core without default-features = false, so today its std feature cannot actually be turned off and the no_std column above still gets a std core one level down. The fix is one line and already open; this PR should land after 0.12.6 is published, with the floor raised.

Doc examples still say std::, deliberately: a doctest compiles as a consumer, and a consumer has one.

What no_std does not mean here

Not bare metal. The core builds for x86_64-unknown-none; concurrent stops at ps-reclaim, whose participant slot lives in thread-local storage and whose no_std path is pthread keys. no_std is a statement about the standard library, not about a class of hardware.

2. The lock-choice benchmark

benches/locks.rs, five commits, already written and never pushed. It prices this crate's lock choice against lock_api+parking_lot, spin, and a real futex, in both the uncontended and the held regimes, with a control so the numbers can be checked. One arm is quarantined rather than deleted, with a note saying why it is broken, and a WFE arm records where that instruction is and is not the right one.

The bench's own copy of upstream parking_lot is renamed parking_lot_upstream: the library's parking_lot is now the fork, one name cannot mean both, and that arm is deliberately the real one with its std futex.

An untracked benches/three.rs was left out. It is a superseded draft of the same measurement and does not compile — it imports arc_swap, which is not a dependency of this crate.

Checks

cargo test --all-features        127 + 1 + 2 + 107 doctests passed, 0 failed
clippy, std and no_std           clean
cargo fmt --check                clean
x86_64-unknown-none              core builds

Version 0.0.13.

Companion to pathscale/congee-wt#6 (same port) and pathscale/parking_lot_lite_hack#3. worktable-vec is what wanted both: pathscale/WorkTable-vec#4 currently declares that its congee and wti features imply std, and these two PRs are what removes that.

meh added 9 commits September 7, 2026 17:42
`concurrent` is the only reason this crate needs `std`, its 31 lock sites are
parking_lot, and `cdc` and `multimap` both imply `concurrent`. So a consumer
wanting only `ChangeEvent` and `Pair` takes parking_lot and libc with them.
DataBucket is exactly that consumer, and it is what blocks its format layer
from being no_std.

Swapping to a spinlock is the obvious answer and the objection is equally well
known: a spinlock burns a core rather than sleeping, which is why parking_lot
exists. Both are true, in different regimes, and nothing here measured which
one this crate is in. parking_lot does not measure it either: its own tests are
two RwLock regression files and assert nothing about parking or CPU.

Measured on 16 cores, medians of 5, with CPU time from getrusage because wall
clock cannot see a burnt core:

  this crate's shape, RwLock<BTreeMap<T, Arc<Mutex<Node>>>>, 200k lookups
    threads    parking_lot cpu     spin cpu     spin vs now
         16          783.1 ms      391.3 ms           -50%
         64          805.6 ms      344.2 ms           -57%

  one lock, held 20000 iterations, oversubscribed
    threads    parking_lot cpu     spin cpu     spin vs now
         16           16.3 ms      116.0 ms          +612%
         32           35.3 ms      291.9 ms          +727%

Short critical sections favour spinning, because parking costs a syscall pair
per contention event and this crate's node access is a few operations on a key
array. A long hold under oversubscription does not, because a spinner holds the
core the lock holder needs to finish, and the wait extends the thing it waits
on.

**Every measurement is one body, generic over `R: RawMutex`**, instantiated
with parking_lot's raw lock and spin's. That is the proposed shape rather than
a benchmarking convenience: both are lock_api underneath, `ArcMutexGuard<R, T>`
is lock_api's type in both cases, and `lock_arc` comes from lock_api rather
than from parking_lot. So `concurrent` can be generic over `R` and let the
consumer choose, instead of naming a lock and deciding for everyone.

The null column is parking_lot against itself and is the floor for believing
any of it. It sits at 1.00x everywhere except the 128-thread contended row,
where it reads 1.60x: that row is not evidence and is left in rather than
quietly dropped.
…claimed

The first version of this bench varied the index RwLock and the node Mutex at
the same time, then reported the difference as a node-lock result. It is not:
with the index lock held constant every node lock lands inside the null.

  this crate's shape, 16 threads, CPU ms
    parking_lot 774.8   spin 782.6   spin+yield 755.5   bounded 807.7   null 1.00x

So the node lock does not matter here and the earlier numbers quoted from this
file, -56% and -50% and 372 against 819, were measuring the index lock. The
index RwLock is the bottleneck and is what deserves the next look.

Where a lock does decide something is a long hold under oversubscription:

  one lock held 20000 iterations, CPU ms
    threads   parking_lot   spin   spin+yield   bounded   futex
        128         153.2  5272.1        987.9     986.9  1680.5

Two things came out of reading parking_lot rather than guessing at it. Its
SpinWait is three exponentially growing pauses, then seven yields, then park,
not the flat 64 spins this file used; copying that schedule made the bounded
arm and the yield arm identical, 986.9 against 987.9, which says the schedule
was never the difference. And the residue is structural: when the schedule runs
out there is nothing to park on without an operating system, so a no_std waiter
keeps running. 6.4x is what that costs.

The futex arm is wrong and is left in only so the next person does not repeat
it. 1680.5 loses to a spinning lock, which a blocking lock cannot legitimately
do. Two rewrites have not fixed it.

Both tables now run the same five locks so a column can be read across them,
and every arm goes through lock_api, which the parking_lot pair shows is free.
WFE idles a core until its event register is set; SEV sets it. ARM documents
this as the intended spinlock construction, and Rust never emits it:
core::hint::spin_loop() on aarch64 is __isb(SY), confirmed by disassembly.
There is no WFE in core::hint and no wrapper in core::arch::aarch64, which is
structural rather than an oversight, because WFE is only useful if the releaser
pairs it with SEV and a one-sided hint cannot express a protocol.

Measured per instruction on this machine: nop 0.3 ns, isb 8.6 ns, wfe 1336.7
ns. So WFE is not a userspace no-op here, and it does not block forever either:
it idles about 1.3 us and returns unprompted.

**And it does not help, which is the point of adding it.** WFE idles the core
but does not yield to the operating system, so a waiter is still a scheduled
thread and the lock holder still cannot get a core. At 128 threads on 16 cores
this arm costs 5584 ms of CPU against 1241 for a yielding spinner. Below the
core count it is level with plain spinning and no better. An earlier version
restarted its spin schedule after each WFE, so it yielded between idles, and
that accident is what briefly made it look competitive.

The rule that falls out: WFE is for threads at most cores, with no operating
system to yield to. That is bare metal and pinned threads, not an
oversubscribed host, and it is written into the arm's doc comment along with
the numbers.

Kept rather than deleted because EKOPathRS is a compiler, and a compiler that
recognises a spin loop can emit WFE where LLVM does not. This measures what
that would be worth and, more usefully, when it would be wrong.
It is not in NAMES and nothing calls it, so clippy was right to flag it. Kept
behind a named module rather than deleted, because the useful thing about it is
that it is wrong: it loses to a spinning lock, which a blocking lock cannot
honestly do, and two rewrites have not fixed it.
The benchmark prices five locks. Nothing checked that the machine underneath it
still behaves the way the numbers assume, so a toolchain or hardware change
would have quietly turned it into a measurement of something else.

Five tests, two in the fast tier. That WFE and SEV execute at all in userspace,
and that a waiter is never stuck when no SEV arrives, since both are what make
a WFE lock safe to write. Behind --ignored: that WFE actually idles rather than
being a nop, that core::hint::spin_loop is not WFE, and the rule itself, that
WFE does not reach the scheduler.

The header carries what was learned rather than only what is asserted: the
instruction and where it exists, that Rust emits __isb(SY) and never WFE and
why that is structural, the per-instruction costs, the rule that falls out, and
the two findings from the same work that should not be re-derived, that
lock_api is free and that the node lock is not this crate's bottleneck.

It also records that the negative result is macOS-specific. Under KVM, WFE
traps and the hypervisor yields the vCPU, which supplies the half that is
missing here, so wfe_does_not_yield_to_the_scheduler is the test to watch on
Graviton. If it starts passing comfortably there, the guidance in this file is
wrong for that host and needs revisiting.
A B-tree over `alloc` collections and an `ftree` that is already no_std. The
only thing here that ever needed an operating system was one `yield_now`.

    ::core::  ops, fmt, slice, hint, borrow, marker, mem, iter, hash, cmp, sync::atomic
    alloc::   Arc, Vec, String, BTreeMap, vec!, format!

The leading `::` is not decoration. This crate has its own module named
`core`, so inside it `use core::x` resolves to `crate::core::x`.

Four things wanted more than a rename:

  - `yield_now` in the two root-publication spin loops falls back to
    `spin_loop` without `std`, which is what the branch above it already does.
  - `published_keys` was a `HashMap<usize, T>`, keyed by a slot number and
    never iterated for order. It is a `BTreeMap` now, the same map without a
    hasher, and the one `alloc` has.
  - `ftree` is no_std, but its `serde` feature takes serde with default
    features, which is `std`. Asking for it unconditionally made every build a
    std build. It now arrives with this crate's `serde` feature, which says
    `std` out loud, as does `superslice-binary-search`: superslice is not a
    no_std crate.
  - `parking_lot` is now `parking_lot_lite_hack`, kept under the
    `parking_lot` name because that is what the code says. This crate uses
    `Mutex`, `RwLock` and `RawMutex`, and the fork has all three. The bench's
    own copy of upstream parking_lot is renamed `parking_lot_upstream`,
    because one name cannot mean both and that arm is deliberately the real
    one, with its std futex.

Doc examples still say `std::`: a doctest compiles as a consumer.

    cargo test --all-features    127 + 1 + 2 + 107 doctests passed, 0 failed
    clippy, std and no_std       clean
    x86_64-unknown-none          the core builds; `concurrent` stops at
                                 ps-reclaim, whose TLS needs an OS
`"^0.1, >=0.1.4"` and `"0.1.4"` describe the same set, and `crossbeam-skiplist`
carried a `"^0.1"` that a bare `"0.1"` already means. One form, used once.
0.12.5 depended on its own core without `default-features = false`, so the
core linked `std` in every build and this crate's no_std column was false one
level down. 0.12.6 is that fix, published today.

    --no-default-features   parking_lot_lite_hack [arc_lock,send_guard]      core []
    --features std          parking_lot_lite_hack [arc_lock,send_guard,std]  core [std]
@pathscale
pathscale merged commit 0c00106 into master Sep 7, 2026
1 of 2 checks passed
@pathscale
pathscale deleted the feat/no-std branch September 7, 2026 11:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant