Load and unload a table as pages, without std, and 0.1.5 - #4
Merged
Merged
Conversation
added 7 commits
September 6, 2026 02:41
The first version invented a third dialect. `LinearTable` in this crate already
says select, len, is_empty, capacity, and WorkTable itself says insert, select,
upsert - and the new table said find_or_claim, find, claimed. A caller moving
between the two tables here had to learn both.
find_or_claim -> upsert returns the existing row or creates one, never
replaces. The difference from a WorkTable upsert
is that the row is then written through &V, since
the value carries its own interior mutability.
find -> select same name and shape as LinearTable::select.
claimed -> len and is_empty beside it, which clippy also wanted.
No behaviour changed and the nine tests are the same nine.
let bytes = table.unload();
let table = LinearTable::load(&bytes)?;
Rows live in a `Vec` while the table is in use and are pages only at rest, so
everything between a load and an unload runs at `Vec` speed because it is a
`Vec`. The format is a codec at the two ends rather than a storage engine
underneath, which is the thing this crate exists not to be.
**No I/O traits.** A load takes `&[u8]` and an unload returns `Vec<u8>`.
Nothing seeks and nothing reads incrementally, so the question of which
read/seek/write abstraction to adopt does not reach this far, and the crate
stays `no_std` with no runtime and no ecosystem attached. Whoever holds the
bytes decides how they got there.
The bytes are DataBucket data pages: a `GeneralHeader` per page, bodies
concatenating into one rkyv archive of the row vector. Not a WorkTable space
file, because there is no schema here to write a `SpaceInfoPage` about, and the
module says so rather than implying compatibility it does not have.
**A fingerprint page, because rkyv will not refuse.** Loading
`Vec<(u64, String)>` bytes as `Vec<(u64, u64)>` first succeeded and returned
correct-looking keys with values like 18446743180901052274, which is a
`String`'s relative pointer read as an integer. Nothing errored. That is the
reinterpretation DataBucket's own migration doc warns about, and validation
cannot catch it because a `(u64, u64)` archive has no invalid bit patterns. So
page 0 now carries a hash of the row type and a mismatch is refused by name.
The hash is FNV-1a over `type_name`: not stable across compilers and not
unique, which is why it fails toward refusing a load rather than toward
accepting one.
The body comes back in an `AlignedVec`. rkyv reads an archive in place and
needs it aligned, and a `Vec<u8>` is aligned to 1; the round trip passing
before this was the allocator being generous, not a guarantee.
The index is derived rather than stored, so a load rebuilds it in one pass
instead of carrying a second thing on disk that can disagree with the rows.
**Cannot merge yet:** `data_bucket_format` is unpublished, so the dependency is
a path into a sibling checkout. It is behind the optional `hydrate` feature,
and the default build is untouched.
The first version of this reached for a page format in another crate, and building that crate was the wrong answer to the question. It is gone. Hydrate is functions in this crate, and it depends on nothing but a serializer. That is also the only option that works. The container `data_bucket` provides is `std`, tokio for file access and eyre through its signatures, and a table that is no_std and alloc-only cannot take that on just to serialize a `Vec`. So the pages are this module's own and deliberately modest: a fixed 24 byte header of little-endian integers, then a body. No rkyv in the header, because rkyv puts an archive's root at the end of its buffer and a header found at a fixed offset should not have to care. Rows use rkyv, where it earns its place. Every page names the row type rather than only the first, so a file spliced onto another is refused at the page where they stop agreeing instead of being concatenated into nonsense. These are not WorkTable space files and the module says so. A WorkTable space opens with a page carrying a name, a schema and a primary key list, and a `Vec<(K, V)>` declares no schema to put there.
table.write(&mut FromStd::new(File::create(path)?))?; // a real file
let table = LinearTable::read(&mut FromStd::new(File::open(path)?))?;
table.append(&mut FromStd::new(appending), from_row)?;
The crate stays no_std. `embedded-io` supplies the two traits, and it already
implements them for `&[u8]` and `Vec<u8>`, so memory needs no adapter and a
`std::fs::File` needs one line of one. There is no second code path for the two
cases and no std feature gating them.
**Every page now stands alone.** It was one rkyv archive split across page
bodies, which meant a single damaged page destroyed every row in the file and
an append rewrote everything. A page now holds an archive of exactly the rows
that fit in it, so damage is one page's problem and appending is writing more
pages onto the end.
**Every header field is checked**, which was not true before: `page` and
`pages` were written and never read, and a field that reads like a guarantee
and is never validated is worse than no field. Gone, replaced by fields that
are: a row count verified against what the body decoded to, a body length
bounded by the page, and a CRC-32.
The checksum is the one rkyv cannot do for us. Its validation says an archive
is structurally sound, which is not the same as saying these are the bytes that
were written: a flipped bit inside a u64 validates perfectly and reads back as
a different number. There is a test that flips one.
Measured on a real file, 200k rows and 12.7 MiB over 814 pages:
flush 67.3 ms 189.0 MiB/s 336 ns/row
open 52.2 ms 243.5 MiB/s 261 ns/row
The first version of the writer took **2343.8 ms**. Filling a page binary
searched over the whole remaining slice, re-serializing every row still to be
written on every probe, for every page: quadratic, and 350x slower than rkyv
encoding the same rows. Now each page starts from the previous page's row count
and walks, so a probe serializes about a page rather than a file. 35x, found by
measuring rather than by reading it.
`read` was wrong. It hydrates a `Vec`, which is a load, and the vocabulary is
the point of the module. So `_to` and `_from` say where and the verb stays the
same everywhere:
unload() unload_to(sink) append_to(sink, first)
load(bytes) load_from(source)
`unload_from(first)` becomes `unload_appending(first)`, because `from` had
started to mean two things in one API: which row, and which source.
`ReadError` becomes `HydrateError` for the same reason.
The module doc was also still advertising `flush(path)` and `open(path)`, an
API that never existed at all.
The workflow installed a target and ran a check against it to prove the crate is no_std. That proves something narrower and less useful: it pins a word size, and a dependency using AtomicU64 fails it for a reason that has nothing to do with no_std. `#![no_std]` enforces itself. A crate carrying the attribute cannot compile a `std::` path on any target, including the host, so the existing `--no-default-features` clippy and test steps already prove it.
I left this half-changed and the crate did not build. That is the first thing this fixes. The header is now DataBucket's `GeneralHeader` byte for byte: seven little-endian u32s in declaration order, which is what `rkyv::to_bytes` of that struct produces, since it has no relative pointers and pads `page_type` from u16 to four bytes. Verified against data_bucket 0.5.7 and pinned by `the_header_is_databuckets_layout`, because a layout reproduced in two places drifts and something has to notice. It is reproduced rather than imported because `data_bucket` is std, through tokio for file access and eyre in its signatures, and this crate is not. **Version 3 is the version that has a row directory.** WorkTable writes 2, and a 2 page carries no directory, which is exactly why a WorkTable data page cannot be read without its index: rows are bump allocated with `data_length` as a high water mark and no delimiters. Eight bytes at the tail of every page fix that here, a row count and a CRC-32, and `every_page_declares_its_own_rows` holds it: the pages account for every row with no index in sight. The count and checksum live in the directory rather than the header because the header is not this crate's to extend. That is the slotted page shape, and it is what makes the format readable by something that did not write it. The `space_id` field carries the row type fingerprint, since a `Vec<(K, V)>` belongs to no space. A WorkTable reader meets a space id it does not know, which is the honest outcome. Unchanged by any of it: flush 67.0 ms, open 52.0 ms over 816 pages.
This was referenced Sep 7, 2026
added 3 commits
September 7, 2026 16:53
The example needed the page size to split a file into pages, and reached for a `hydrate_page_size()` that does not exist. `PAGE_SIZE` is already a `pub const`; the module is private, so nothing outside could see it. Export it rather than adding a function that returns it. Page decode is the only embarrassingly parallel part of a load, so this is the measurement that says whether threading it is worth anything at all.
Enabling `arctic` enabled arctic's default features, which are `std` plus an
SMR backend, so a crate that declares `#![no_std]` quietly linked one anyway.
arctic 0.1.11 is the first release where this is fixable: before it,
`smr-ps-reclaim` forced `std` on by itself.
--features arctic arctic-wt [smr-ps-reclaim] ps-reclaim [libc,spin]
--features arctic,std arctic-wt [smr-ps-reclaim,std] ps-reclaim [libc,spin,std]
`congee` and `wti` now state `std` in the manifest, because neither crate is
no_std upstream and pretending otherwise only moves the failure later.
CI grows three steps in the existing job, no new runner:
- the core and `hydrate` build for x86_64-unknown-none, which has no `std`
to find, so the no_std claim is checked against a target rather than a host
- the arctic graph is asserted to carry no `std` feature; the assertion was
run against `--features arctic,std` first to prove it can fail
The arctic backend itself cannot target bare metal: ps-reclaim stores its
participant slot in thread-local storage and its no_std path is pthread keys,
so it needs an OS. no_std here means no standard library, not no operating
system, and that distinction is now written down instead of assumed.
added 4 commits
September 7, 2026 17:06
0.0.12 replaced the skiplist topology with ordered routing, took the shared reader bottleneck out of lookups, and hardened publication invariants. It is on WTI master and published, and WorkTable already resolves it. This crate was the last consumer still floored at 0.0.11, so it was the only one running the old topology.
The doc said keys are `u64` and that `usize` was chosen because `AtomicU64` does not exist on 32-bit bare-metal targets such as `thumbv7em-none-eabi`. Both halves are wrong. The API takes `key: usize` and always has, and nothing here is built or tested for a 32-bit target, so the justification cited support that does not exist to explain a choice that needs no excuse: the key is the same width as the slot arithmetic that indexes it. no_std is a statement about the standard library, not about a class of hardware. The wording no longer implies otherwise.
`scatter` kept a second golden-ratio constant for a 32-bit `usize`, arriving with the same commit as the `thumbv7em-none-eabi` claim in the doc above it. Nothing here is built or tested for a narrow target, so that branch was code written for support that does not exist. It cannot simply be deleted: truncating the 64-bit constant leaves `0x7F4A_7C15`, which is even, so the multiply stops being invertible and keys quietly collapse onto the same slot. Silent wrong hashing is a worse answer than no support, so the crate now says which it is at compile time, the way WorkTable's codegen already does for a `u64` congee key. Verified both ways: a 32-bit target stops with "worktable-vec's AtomicKeyTable requires a 64-bit target", and x86_64-unknown-none still builds clean.
…l not load
`unload()` returned `Vec<u8>` and could not fail. A row whose archive did not
fit a page body got a page of its own anyway, spilling past the page boundary;
`load` then stopped at `Overlong` and every row in the file was unreachable.
The write reported success. Measured before the fix: one 20 KB row wrote 32,768
bytes and lost the row, and a 30 KB row among a hundred ordinary ones lost all
hundred and one.
The limit is 16,348 bytes of archive per row, which is a page minus the header
and the directory, and it is a limit rather than a spill because a page holding
part of a row stops standing alone. That property is the reason the format
exists.
unload, unload_appending -> Result<Vec<u8>, RowTooLarge>
unload_to, append_to -> Result<(), UnloadError<W::Error>>
`RowTooLarge` names the row, the size it needed and the size available.
`UnloadError` mirrors `HydrateError`, splitting a sink that would not take the
bytes from rows that cannot be written at all.
Four tests, and all four fail with the check disabled: the refusal itself, the
row it names, the boundary asserted from both sides, and a refused write
leaving the sink untouched.
cargo test --all-features 35 passed, 0 failed
clippy, fmt, x86_64-unknown-none clean
This was referenced Sep 7, 2026
added 2 commits
September 7, 2026 18:07
`"^0.1, >=0.1.11"` and `"^0.1.11"` describe the same set, and the second one
says it once. A bare `"0.1.11"` is already a caret, which is the form the rest
of the house uses.
`wti` is the one that changes meaning. Caret on a `0.0.x` version pins the
last number, so `"0.0.12"` will not take 0.0.13 the way `"^0.0, >=0.0.12"`
would have. That is what a caret means for a crate below 0.1.0 - every release
is breaking - so the bump becomes deliberate rather than automatic.
cargo test --all-features 35 passed, 0 failed
Both crates were std-only, so this crate declared it rather than pretending
otherwise. congee-wt 0.4.5 and WorkTablesIndex 0.0.13 changed that, so the
declaration is now wrong and comes out.
--features arctic arctic-wt [smr-ps-reclaim] ps-reclaim [libc,spin]
--features congee congee-wt [] ps-reclaim [libc,spin]
--features wti WorkTablesIndex [concurrent]
parking_lot_lite_hack [arc_lock,send_guard]
ps-reclaim [libc,spin]
with std added every one of them gains it, and nothing gains it
without asking
WorkTablesIndex moves by hand rather than by caret: `0.0.12` will not take
0.0.13, because a caret on a `0.0.x` version pins the last number. That is
what a caret means below 0.1.0, so the bump is deliberate.
The CI assertion covered arctic alone and now covers all three, together and
separately. Run against `--features arctic,congee,wti,std` first, where it
fails, because a check that cannot fail proves nothing.
all features 35 passed
no_std + arctic / congee / wti 10 passed each
x86_64-unknown-none core and hydrate build
clippy, fmt clean
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One PR, one repo: the atomic-table vocabulary, the whole
hydrateinterface, and theno_stddecoupling, rebased into a single linear branch. Ten commits, zero merge commits.1. Vocabulary
AtomicKeyTablegains the names the rest of WorkTable uses, so a reader moving between them is not learning a second dialect.2. Pages, and a disk that is a trait
A table you cannot put on a disk is a benchmark, not a component. The whole persistence interface is two operations —
unload()to pages,load()back.PAGE_SIZEis 16 KiB, each page stands alone, and a page carries a row directory (format version 3, matching DataBucket) so a reader can find row boundaries without decoding the page first.hydratetakesembedded-io, notstd::fs. A caller with an OS wraps aFilethroughembedded-io-adapters; a caller without one plugs in whatever it has. Only the tests linkstd, so they can put a page run on a real disk and show the traits reach one.rkyvandembedded-ioare optional and pulled in only byhydrate, because serialization is the one thing here that needs a serializer.Where an open actually spends its time
examples/where_time_goes.rs, 200,000 rows, 816 pages, median of 5, release:Decode is not part of a load, it is the load — the two agree inside noise. Page decode is embarrassingly parallel and everything around it is free, so threading it is worth roughly the full core count. Nobody had to guess.
The example called a
worktable_vec::hydrate_page_size()that never existed.PAGE_SIZEwas alreadypub constbehind a private module, so it is re-exported rather than wrapped in a function.3. The
no_stdclaim, made trueThis crate declares
#![no_std]. Enablingarcticenabled arctic's default features —stdplus an SMR backend — so it linked a standard library it says it does not use. It went unnoticed because it was only ever built on a host that has one.arctic 0.1.11 is the first release where this is fixable: until pathscale/arctic-wt#9 landed,
smr-ps-reclaimforcedstdon by itself, sono_std+ reclamation was not a combination that existed.A
stdfeature now exists and forwards; it is off by default, because the core never needed one.congeeandwtideclare"std", because neither isno_stdupstream —congee-wthas no#![no_std]at all andWorkTablesIndex'sconcurrentfeature pullsparking_lot.What is no_std, precisely
x86_64-unknown-nonehydratearcticcongee,wtiThe arctic row is not a defect to fix here. ps-reclaim keeps its participant slot in thread-local storage and its
no_stdpath is pthread keys.no_stdmeans no standard library, not no operating system, and that is now written down rather than assumed.The verifier
A rule with no check is a comment. CI grows three steps in the existing job — no new runner:
hydratebuild forx86_64-unknown-none, a target with nostdto findstdfeatureThe assertion was run against
--features arctic,stdfirst and it failed there, which is the only evidence it can fail at all.Checks
Version 0.1.5.
no_stdis also no longer gated on a CPU in CI — that was a target property being asserted about a runner.