Skip to content

Columnar readers keep every column, a stable Arrow type, and exact DECIMALs (#113, #114, #115) - #132

Merged
tae898 merged 1 commit into
mainfrom
fix-columns
Oct 3, 2026
Merged

tae898 merged 1 commit into
mainfrom
fix-columns

Conversation

@tae898

@tae898 tae898 commented Oct 3, 2026

Copy link
Copy Markdown

Fixes #113, #114, and #115 (the columnar readers to_columns, to_dataframe, and to_arrow, and the header of export_to_csv). Like #131 this changes the Java bridge (ColumnBatcher), so the tests need the bridge jar built from this branch, which the wheel build does.

#113, a property the first row lacks was dropped. ColumnBatcher took the column set from the first row, and the Python side pinned the first batch's set for every later batch. Each batch now reports the union of its rows' property names in order of first appearance, and the Python side takes the union over the batches in order of first appearance, filling a batch that lacks a column with nulls in the shape that column has elsewhere (NaN, NaT, or None; pa.nulls for Arrow). The three readers return n, kind, reason for the issue's three events at every batch size. export_to_csv takes its header from the union of the first batch's rows (and of every row for a list of dicts); a column that first appears after the first batch of 10,000 rows now raises ArcadeDBError naming it and saying to pass fieldnames, instead of the writer's bare "dict contains fields not in fieldnames" (the header is already written by then, so the partial file remains).

The union costs one pass over each row's property names, because Result.getPropertyNames() builds a set per row. On a 200,000-row, twelve-property to_columns() I measured about 25% (medians 0.91 to 1.01 s on the old bridge jar against 1.15 to 1.24 s on the new, three alternating rounds pinned to cores 0 to 3 of the laptop; relative only, the laptop was running other work). So to_columns() and to_arrow() take an optional columns=[...] that reads exactly those columns as a projection would and skips the pass; a row lacking one reads null and a property not listed is left out.

#114, an Arrow column's type followed the batch size. A chunk that says nothing about its column's type (every value null, or a list<null> from only empty lists) now takes the type of the other chunks instead of forcing the whole column to strings; level is int64 and tags is list<string> at batch sizes 25,000, 2, and 1 for the issue's five readings. A column that mixes types inside one batch (an int in one row, a string in the next) raised a raw ArrowInvalid; it is now the string column the cross-batch case already produced.

#115, a DECIMAL column came back as JSON numbers. BigDecimal gets its own column type in ColumnBatcher (the string layout, each value written with toPlainString), decoded to Decimal objects: an object array for to_columns() and to_dataframe(), a decimal128 Arrow column (decimal256 above 38 digits, an exact string above 76), with differing precision and scale across batches unified to one decimal type. The issue's four cases (36 digits, whole amounts only, whole and fractional, a whole amount above 2**63) plus one with a null round-trip exactly at batch sizes 25,000 and 1.

The tests are in tests/test_columnar_readers.py (31) and are described in the testing docs. Verified: 29 of the 31 fail on the bridge jar and Python of main (the two that pass are 25,000-row cases that already worked) and all pass on this branch, the existing arrow, resultset, and exporter tests pass, and the full suite passes (540 passed, 3 skipped, 3 xfailed), bandit clean at CI strictness.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL

…rrow type, and exact DECIMALs

to_columns, to_dataframe, and to_arrow took the column set from the first row
and pinned the first batch's set for the rest, so a property the first row
lacked was dropped. Each batch now reports the union of its rows' names and the
readers take the union over the batches in order of first appearance, with
nulls for a batch that lacks a column. Finding the names costs one pass over
each row's names (about 25% on a 200k x 12 to_columns), so to_columns and
to_arrow take an optional columns= that reads exactly those columns. The CSV
export's header is the union of the first batch, and a column that first
appears later raises an ArcadeDBError that names it.

An Arrow column's type no longer follows the batch size: a chunk that is all
null or holds only empty lists takes the type of the other chunks, and a column
that mixes types inside one batch becomes strings instead of raising
ArrowInvalid.

BigDecimal has its own column type in ColumnBatcher, decoded to Decimal objects
(an object array) and to a decimal Arrow column, so no digit is lost to a
double, the dtype no longer follows the data, and a value above 2**63 no longer
overflows.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL
@tae898
tae898 merged commit 8c0aeb6 into main Oct 3, 2026
48 checks passed
@tae898
tae898 deleted the fix-columns branch October 3, 2026 15:30
tae898 added a commit that referenced this pull request Oct 6, 2026
…ith a columnar insert in the bindings (#211) (#212)

* l2 (re-pin prep, NOT for October): both ArcadeDB graph arms load through the bulk path (#8287; BUGS F123)

Embedded: graph_batch(use_wal=True, expected_edge_count) for persons+KNOWS, graph_batch(use_wal=True)
for the message half with endpoints from returned RIDs. Served: POST /api/v1/batch?wal=true
(&expectedEdgeCount for KNOWS; idMapping=true so the message request references Persons by RID),
streamed JSONL. Laptop micro smoke on 5b040a438e: all 20 answer digests/counts identical to the
current loaders; build 1.75 -> 1.03 s embedded, 7.38 -> 1.21 s served. Message half not smoked
(needs the LDBC corpus): smoke on mini after ALL-DONE.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* l1tpc (re-pin prep, NOT for October): the served loader binds each batch as INSERT ... CONTENT :rows (#8337, DECISIONS #116 item 4)

Replaces a sqlscript of 2,000 INSERT ... SET statements with the values in the
text. The batch size is BENCH_SERVED_LOAD_BATCH (default 2000), recorded in the
row as served_load_batch, for the 2k/5k/10k sweep on the bench host.

Laptop smoke, SF0.1, served arm on 7effa0950e, main's loader against this one:
600,572 line items both, all five answer digests identical, ingest 52.1 s -> 43.4 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* l3d + page (re-pin prep, NOT for October): the served dense build binds INSERT ... CONTENT :rows; ingest sentences follow (#8337, DECISIONS #116 item 4)

Replaces 500-statement sqlscript batches with each vector spelled out as text.
From this Python harness CONTENT is the fastest served vector path (laptop, 50k x
96, batch 2,000: sqlscript literals 2,754 rows/s, psycopg over the Postgres wire
3,990, CONTENT 4,980; every path stores the float32 values exactly), so it is used
instead of the Postgres wire #8337 found fastest from Java. Batch size is
ArcadeServer.load_batch (BENCH_SERVED_LOAD_BATCH, default 2000).

The October ingest sentences for l3d and the documents tables now describe the
bound batches, pinned to the load_batch attributes; SERVER_BATCH is retired.
Smoke, micro, served arm on 7effa0950e: recall@10 1.0 both, ingest 2.71 -> 1.36 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* l3d (re-pin prep, NOT for October): the served dense search binds its query vector, index name, k and ef (DECISIONS #116 item 2)

As the embedded arm already does. Laptop probe, 20k SIFT, 300 queries: identical
top-10 on all 300, p50 5.40 -> 4.37 ms (repros/vector-query/served_bound_vector_probe.py).
Micro smoke with the index name bound too: recall@10 1.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* l1tpc (re-pin prep, NOT for October): ArcadeDB served and SurrealDB OLTP statements bind their values (DECISIONS #116 item 2)

ArcadeServerTPC: new-order and payment stay one sqlscript request each, with
named parameters shared across the script's statements; the four CRUD
operations bind :c/:pk. SurrealTPC (and the served twin, which inherits it):
$vars with RecordID objects, so record-id addressing is kept and nothing is
written into the SurrealQL text.

Laptop smoke at SF0.01 (BENCH_CRUD_OPS=500), main against this branch:
answer digests identical for arcadedb_server and surrealdb_tpc. Timing is
sub-ms on the laptop and not reliable there; mini measures at the re-pin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* l2 (re-pin prep, NOT for October): the graph reads and writes bind their values on every engine that can (DECISIONS #116 item 2; BUGS F125, F130)

The shared Cypher template takes $id, $new_id and $name, passed through each
engine's own driver: ArcadeDB embedded (query/command with a params map) and
served (HTTP params), Neo4j and Memgraph (session.run params), FalkorDB
(query params), LadybugDB (prepared once per text per connection, then
executed with params). SurrealDB binds $vars with RecordID objects, and its
write is now one BEGIN/COMMIT transaction: it was two statements in one
request, which SurrealDB runs as two transactions (F130). DuckPGQ keeps its
values in the text and says so: it rejects any parameter in a GRAPH_TABLE
query (cwida/duckpgq-extension#75); its writes and scan already bind.

Laptop micro smoke, 9 engines, October head against this branch: all 8
answer digests identical on every engine. Timings are laptop numbers, for
direction only: Neo4j point 30.2 -> 7.2 ms, ArcadeDB embedded 2.17 -> 0.32,
served 5.9 -> 3.5, Memgraph/FalkorDB/LadybugDB about level, SurrealDB served
write 14.3 -> 7.6 (one transaction). SurrealDB embedded hop1 read 5.1 -> 9.8
under CPU contention from a probe; being checked with an in-process A/B.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* e2, l3d, l3s + page (re-pin prep, NOT for October): bind the remaining timed values; one set-based e2 update; served sparse and mutate loads through CONTENT (DECISIONS #116 items 2 and 4; BUGS F130, F131)

e2: both ArcadeDB arms bind pid and the id lists (`pid IN :ids` still reads
the unique index, EXPLAIN: FETCH FROM INDEX) and send ONE set-based update
instead of one statement per product (F131; served: one request instead of up
to six); SurrealDB binds $q and RecordID lists and updates with one
`UPDATE $ids`; PG+AGE's single-pid read binds through cypher()'s third
argument. l3d: the deletes bind (ArcadeDB `IN :ids`, DuckDB `IN (SELECT
unnest(?))`), Milvus deletes by ids=, SurrealDB binds its query vector and
deletes one `DELETE $ids` per batch (it was 200 statements, 200 transactions:
F130); LanceDB's predicate-only API is declared. The served mutate insert and
the served sparse build load through `INSERT ... CONTENT :rows`. Page: the
graph ingest sentence now describes item 1's bulk path, the sparse one the
bound batches, each pinned.

Laptop smokes, before/after, answer digests identical everywhere they exist:
ArcadeDB served arms on the upstream-main pair (7effa0950e; a first pass
without the image export ran 26.8.1 and is discarded): graph point 5.75 ->
3.30 ms, e2 transaction 26.7 -> 14.7 ms, dense recall 1.0 with mutate insert
0.655 -> 0.265 ms/vector, sparse recall 1.0. Comparators: every e2 digest and
atomicity result unchanged; dense recall and mutate results unchanged. One
regression, fixed in the next commit: PG+AGE loses its pid expression index
for a bound LIST.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* e2 (re-pin prep): PG+AGE keeps its pid LIST in the text; a bound list loses the expression index (2.0 -> 26.4 ms), the single pid stays bound (DECISIONS #116 item 2)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* e2 (re-pin prep): SurrealDB's record-id ranking binds its vector and candidates, on top of main's F132 fix (rebase onto #117)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* l3d (re-pin prep, NOT for October): the embedded dense ingest goes through graph_batch with the WAL on, one Java float[] per vector (DECISIONS #116 item 3)

Replaces a SQL INSERT per vector in 10,000-row transactions with the engine's
bulk loader, WAL kept on as the maintainers recommend (#8287), fed 10,000
vectors per commit so the live Java arrays stay bounded at deep10m. The
October ingest sentence says so, pinned to BATCH.

Laptop, pinned pair 7effa0950e, the tier's 20,000 SIFT vectors: ingest 1.84 ->
1.46 s fp32 and 1.82 -> 1.49 s int8, recall 0.9996 -> 0.9998 and 0.9953 ->
0.9951 (HNSW build noise), query p50 level. The 1.4x of the 2026-09-24 audit
was measured at 200k; mini measures it at the re-pin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* page (re-pin prep, NOT for October): the heap setting printed beside the peak memory for JVM engines (DECISIONS #116 item 7)

A `JVM heap` column right after `peak memory GiB` on every table that has
one, from the rows behind each cell at its own size: ArcadeDB, Neo4j (and
the composed stack's Neo4j half), Elasticsearch, and QuestDB record a heap;
every other engine prints a dash. A JVM engine's peak sits close to the heap
it was given, so the decision keeps the peak (what an operator provisions)
and puts the setting next to it; no new instrumentation. The lookup is shared
with the memory note (_entry_heaps), so the two cannot name different heaps.

Dry landing from this branch (e2 scoped, October rows): every gate passes;
l2 shows 4g/12g beside ArcadeDB's and Neo4j's peaks, e2 6g/16g.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* e2 + page (re-pin prep, NOT for October): the served cross-model build loads through /batch with the WAL on (DECISIONS #116 items 1 and 4)

Products and RELATED edges stream as JSONL to POST /api/v1/batch
(?wal=true&expectedEdgeCount=N), each embedding as the float32 values'
exact doubles, as the graph lane's served arm loads (#8287); the counts the
server returns are checked. It replaces CREATE VERTEX / CREATE EDGE ... FROM
(SELECT ...) statements with every value pasted into sqlscript, where the
embeddings crossed at six significant digits. The October e2 ingest sentence
says so; September's keeps what September ran.

Laptop, 7effa0950e pair, 50k products: build 40.5 -> 17.4 s (ingest 28.7 ->
9.2 s), all three answer digests equal the embedded arm's, retrieval recall
0.785 -> 0.7855, atomicity 0 torn of 40 unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* e2 PG+AGE: the pasted pid list cites apache/age#2582

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL

* e2 (re-pin prep, NOT for October): the embedded arm's ranking statement takes the query vector as a Java float[] (BUGS F145)

_vec_topk and hybrid_op already pass it that way; _rank_candidates passed a
Python list, which crosses JPype one element at a time. Laptop, 50,000
products, 39 candidates: 1.42 -> 1.00 ms p50 with the October statement,
0.80 ms with the ids bound too; the same top ten on 300 of 300 queries.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77

* graph lane (re-pin prep, NOT for October): real LDBC ages, a filter that filters, a per-read budget, and a digest that survives six-digit ties (BUGS F146, DECISIONS #119, #120)

Ages: ldbc_snb read the corpus's epoch-millisecond birthday as a date
string, so every age was 0 and hop3f's x.age > 30 matched nobody. Parsed
as milliseconds (ISO kept as the fallback), LDBC ages span 36-46, so the
threshold moves to 41 (keeps 49.7% at SF1, 49.3% at SF10), one constant,
graph_common.HOP3F_MIN_AGE, read by all five spellings.

Per-read budget (#120): each transactional read gets one budget across
both passes (budget_lookup, lane default OLTP_READ_BUDGET_S = 1,800 s,
clamped to 540 s at SF1). A read that spends it stops, keeps percentiles
over the starts it answered, records <read>_censored/_iters/_budget_s,
and declares its answer with bench_common.record_censored_answer; the
gate lists it under E4b, the page says it is not compared (never
'cannot express') and names it in the per-engine budget sentence, now
worded right for a single cell. Found because SurrealDB embedded's 2.3.10
core dedups quadratically and its SF1 cell timed out whole.

Digest: floats round twice (12 significant digits, then 6), because exact
six-digit ties (41.15625, a mean of 32 ages) split Neo4j's running mean
from eight engines on hop1. Tie and declaration tests added.

Laptop smoke, SF1, ten engines at this tree: 0 equivalence failures, 8
groups agreeing (hop3f on real, non-zero answers), SurrealDB embedded's
hop3f censored after 111 of 500 starts with its other reads and writes
measured, the cell done in 996 s. PAGE-SPEC and the fairness comment
updated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77

* l4 (re-pin prep, NOT for October): ArcadeDB's native time-series arms settle until the engine reports no uncompacted sample (DECISIONS #121)

settle() polls SELECT FROM schema:types until every shard's mutableSamples
is 0 (bounded, BENCH_TS_COMPACT_WAIT_S, default 300 s), outside the ingest
timer, as QuestDB's arm waits for its WAL; the driver records
ts_mutable_at_query and ts_compaction_wait_s at the first timed query, so
a row proves its compaction state. Why: ArcadeData/arcadedb#8574, the
newest reading is 21 ms on the mutable tail and 0.13 ms compacted.

Also removes an asymmetry: the embedded native arm slept the lane's
BENCH_TS_SETTLE_S itself and the driver then slept it again (180 s at
October's 90 s) while the served arm's settle was empty. The symmetric
floor stays the driver's. The shard list is read as a map or its JSON
text (the October wheel's to_json_list renders a nested result as a
string); anything unreadable counts as -1, never 0.

Laptop smoke at ts100, BENCH_TS_SETTLE_S=0: served arm waited 55.5 s,
embedded 5.1 s, both ts_mutable_at_query 0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77

* runner + l3s (re-pin prep, NOT for October): a server that fails to start is recorded, a raising cell no longer ends its batch (F150); the served sparse rows are the generator's lists (F151)

F150: run_cell's finally reads _therm0, which is taken only once the client
starts, so since dea7115710 every early return before that (server_not_ready,
the server-heap mismatch) raised UnboundLocalError instead of returning its
row. And worker() had no except: the exception killed the thread, the cells
still pending on it never ran, and with no error row the runner exited 0.
Now _therm0 is initialised beside cli_cid, and a raising cell is printed as
RAISED <cell>, the batch continues, and the exit is 1. Not fired on mini in
October (no crash in STATUS.txt). Laptop: the JDK 21 served cell that
crashed now records server_not_ready and exits 1.

F151: item 4's served sparse loader passed every weight through
np.asarray(..., float32) and back although both corpus sources already yield
Python lists (rows compare equal without it): 5.30 s of client CPU at 100k
on BigANN, 0.08 s without. Laptop, BigANN 100k, c44f6604e8 pair: served
build 56.3 / 62.0 s against October's loader 63.2 / 70.0 s (embedded
23.6 / 26.4 s); the re-pin loader stores every sampled weight exactly
(October's 9-decimal text does not, 80 of 101 records off by up to 3.4e-5).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77

* l1tpc (re-pin prep, NOT for October, NOT yet adopted): #8478's bulk path for LineItem, measured, with a bucket-count control (CAMPAIGN item 10, ingest)

The answer on ArcadeData/arcadedb#8478: async createRecord on a type with
buckets = k x the executor's parallel level (k 1 or 2), set at CREATE TYPE.
Applied to the embedded documents loader: LineItem created with BUCKETS =
writers x BENCH_ARCADE_BUCKETS_PER_WRITER (default 1), loaded through
insert_many(parallel=True) with writers x commitEvery rows per call
(waitCompletion() commits every writer's open batch, so smaller calls mean
smaller commits), the executor's WAL flush set per class (its writers stamp
their own and ignore txWalFlush: yes_full at strict, via
bench_common.arcade_async_sync), the stored count checked after the load (an
older wheel drops rejected records silently, F153), and the row records
lineitem_buckets, async_writers, async_sync and load_call_rows.
BENCH_ARCADE_LINEITEM_BUCKETS overrides the count for control runs; both
variables are on the runner's container allowlist.

Laptop, SF0.1, cores 0-3 (3 writers), engine c44f6604e8: ingest 1.6x faster
at both classes, every answer digest identical. But a control on this same
harness with only the bucket count varied puts the five analytics queries
17-47% slower at 3 buckets than at 1. That contradicts the answer's "no
longer costs on scans", so this is prepared, not adopted: the question goes
back upstream with a Java repro before the re-pin decides.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77

* l2 (re-pin prep, NOT for October): PostgreSQL + AGE joins the graph tables (DECISIONS #128, #131)

A pgage_graph adapter on dbbench:pg-age: the lane's own Cypher through
AGE's cypher(), values bound as one agtype map, rows parsed back from
agtype after the timed call. Loaded by COPY into AGE's label tables with
the graphids AGE's sequences would assign (20,000 persons in 0.12 s and
418,599 KNOWS in 2.9 s on the laptop); Message is a PostgreSQL
inheritance parent of Post and Comment, so LSQB's (:Message) finds both.
One added index, a btree on the Person id property, by measurement
(#112; point 58x, hop1 16x, hop2 3.2x, hop3f 2.1x, same answers).

Smoked on the laptop: micro oltp and olap, the LDBC SF1 slice, and the
capped sf1full slice; every digest matches Neo4j's and LadybugDB's except
the two CRUD read-backs at micro, where AGE's btree scan starting at a
>= bound skips the equal key (apache/age#2587, filed today, every AGE
release since 1.6.0). equivalence_check declares it as KNOWN.

Also: Dockerfile.pgage pins postgresql-18-age=1.8.0~rc0-2.pgdg12+1 (the
build the cross-model rows carry; CAMPAIGN section 7 row 26), and both
laptop smoke scripts run their LSQB stage at sf1full: at sf1 the message
half never loads, so no smoke since 2026-09-19 asked LSQB's nine.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* runner + l2 (re-pin prep, NOT for October): PostgreSQL + AGE's query pools fitted to the cell like every other engine's (FAIRNESS F3/F6)

The user: "make its resources usage (cpu, memory, heap, etc) fair as we
do the other dbs". Both AGE arms (pgage_graph, pg_age_e2) now get
parallel workers per query = cpuset - 1 (the leader is the last core),
max_parallel_workers = cpuset, and work_mem = (cap x 0.75) / (cpuset x
16), beside the cap-sized shared_buffers, effective_cache_size and
maintenance_work_mem they already had. PostgreSQL's fixed defaults (2
workers, 4 MB) reached 3 cores of 12 and spilled instead of using the
cap, where every other engine's pools reach the cell.

Measured on the full LDBC SF1 network (3.16M message vertices, 13.6M
edges), 8 cores and a 16 GiB cap, repros/age-dialect/
resource_fit_probe.py: workers fitted made LSQB q2 1.8x faster and q4,
q5, q7 1.1-1.25x; the fitted 96 MB work_mem cut temp-file spill from
42 GB to 0.7 GB per analytics pass at no net time cost; 384 MB bought
nothing more; no OOM kill; answers identical across all four settings.
The graph adapter reads every setting back onto the row (pg_*), the
cross-model rows carry them in server_cmd. Smoked through the runner:
l2 micro oltp+olap (digests unchanged) and e2 hybrid.

FAIRNESS F6 gains the PG+AGE row; COMPARATORS names the graph arm and
the pinned package.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* l3d (re-pin prep): Elasticsearch, Memgraph, FalkorDB, and LadybugDB join the dense tables; DuckDB VSS and ArangoDB split ingest from index (DECISIONS #131 item 3, #132)

Five new arms, each at the matched operating point in its own unit and
with its pools fitted to the cell and read back onto the row:
elasticsearch_dense (hnsw) and elasticsearch_dense_int8 (int8_hnsw, top
k re-scored), on the sparse arm's 9.5.4 image and heap; memgraph_dense
(3.13.1, the graph arm's image and flags; its USearch config is fixed at
connectivity 16 / expansion_add 128 and any key we pass is ignored);
falkordb_dense (6.0.1, since 4.x crashes on a write in a vector
statement; the background-built index is waited for to OPERATIONAL,
recall 0.05 without the wait); ladybug_dense (the official vector
extension, ml = 2M base degree, COPY from Arrow). DuckDB VSS and ArangoDB
now time their load and their index separately, so
PHASE_SPLIT_DISCLOSED is empty and the two Elasticsearch arms are
declared (F14c). l3d_dense and dense_multipass_driver record each
adapter's row_extra.

Smoked through the runner at micro (mutation forced on) and tiny: recall
0.98-0.9997 beside Qdrant's 0.9999, deleted hits 0, F14c ok, base degree
32 everywhere. Two engine defects found by the mutation phase: LadybugDB
0.20.4 crashes when searching after re-inserting deleted keys (0.21.1,
the re-pin's, is clean); Memgraph keeps committed deletes in its vector
index until GC (memgraph/memgraph#4975), so its delete runs FREE MEMORY
inside the timer. NOTES-dense-arms-20261002.md has each arm's status and
the decisions left.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* l2 graph: MongoDB's two- and three-hop reads use $graphLookup (DECISIONS #131 item 2, CAMPAIGN rows 27 and 28)

hop2 seeds $graphLookup at the person with maxDepth 1 and keeps the depth-1 edges; hop3f and the three-hop visited probe seed it at each first edge's end and keep the depth-1 edges that are not the first edge. Both are exact when the graph has no self-loops, which build() now counts onto the row (mongodb_knows_self_loops) and refuses. A probe found 0 disagreements with the chained $lookup form over 200 ids on micro and the LDBC SF1 slice; seeding three hops at the person disagreed on 92 and 12 of 200. hop1, the writes, and the analytics stay $lookup, and QUERY_LANGUAGE says so.

Found by the capped SF1 slice smoke: run_write inserted the new person and edge even when the anchor person was not loaded, where the Cypher engines' MATCH creates nothing, so the write and update digests differed. It now looks the anchor up first, in the same transaction. main() also reads row_extra after the build.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* e2: Memgraph, LadybugDB, and DuckDB join the cross-model lane (DECISIONS #131 item 4, CAMPAIGN row 39)

Each runs vector top-k, the hop, and the update in one transaction, and the atomicity probe by rollback, like the existing e2 arms. memgraph_e2: served on the graph arm's image and pools, the vector index at its defaults (Memgraph takes no HNSW parameters; disclosed), search breadth 100 then keep k (asking for k gave recall 0.584). ladybug_e2: embedded, the official vector extension, HNSW matched to the table (mu 16, ml 32, efc 100, efs 100), threads and buffer pool fitted (F160), rollback on any exception. duckdb_e2: tested and added, vss HNSW (M 16, ef_construction 100, ef_search 100) under the experimental persistence flag, a DuckPGQ hop with the pid list inlined (DuckPGQ binds no parameter), threads fitted from the affinity mask and memory_limit from the cgroup. Durability is read back for all three. e2_hybrid.main() records each adapter's row_extra.

Laptop smoke at e2 scale: all four graph-capable arms (these three and pg_age_e2) agree on every digest; atomicity 0 torn of 40 for each. Registered in runner (BACKENDS, LANES e2), export_web display names, page_check (optional until measured), and fairness_check (e2 index map; LadybugDB and DuckDB strict-only). COMPARATORS and FAIRNESS F6 rows name the arms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* export_web: PostgreSQL + AGE joins the multi-model table (DECISIONS #131 item 9, CAMPAIGN row 43)

MULTIMODEL_ENGINES gains "PostgreSQL + AGE". MULTIMODEL_ALIASES folds the cross-model arm's name ("PostgreSQL + pgvector + AGE") into that row, because both arms run the dbbench:pg-age image; PostgreSQL's documents, pgvector, and TimescaleDB arms run other images and stay their own engines. page_check's roster uses the same fold. PAGE-SPEC says five engines.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* l4 + runner (re-pin prep): plain PostgreSQL on time series, the PostgreSQL pool fit on every arm but the defaults arm, the quantization survey (DECISIONS #131 items 6 and 8, FAIRNESS F6)

postgres_ts: COPY into a plain table, the (host, ts DESC) btree TimescaleDB
carries (measured at ts100: q_last 108.95 -> 0.48 ms, q_range 92.91 ->
1.08 ms; the 12 h aggregate pays 136 -> 217 ms, declared), date_trunc
buckets in UTC; smoke digests equal TimescaleDB's and SQLite's on all six.

runner.PG_FIT_CMD factors out the AGE fit and applies it to
postgres_tuned (its constant 64MB work_mem removed), timescaledb,
pgvector_dense, pgvector_sparse, and postgres_ts; not to `postgres`, the
defaults arm. BENCH_PG_FIT=off formats PostgreSQL's defaults for a
same-run A/B: answers identical in every pair, no OOM kill (FAIRNESS F6
row). The time-series PostgreSQL arms read every setting back (pg_*).

QUANTIZATION.md: what each dense and sparse comparator ships against our
DDL. Neo4j 2026.08.1 defaults its vector index to BINARY, so the dense and
cross-model Neo4j arms record fp32 while running binary (verified on the
pinned image); Elasticsearch 9.5.4 defaults to int8_hnsw at our dims;
LanceDB 0.39.0 does build an unquantized IVF_HNSW_FLAT.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* l5 (re-pin prep): SQLite and DuckDB embedded on the lifecycle table (DECISIONS #131 item 5, queued last by #133)

l5_lifecycle_sql.py runs the lane's situations, sizes, session, and mode
set on SQLite and DuckDB in-process, streaming the build in batches
(F161); what a model lacks is declared with the engine's own error.
DuckDB builds vector (VSS HNSW) and graph (DuckPGQ property graph, read
with GRAPH_TABLE); SQLite builds the document and time-series situations.
Read digests equal ArcadeDB's and SurrealDB's at lc10k (doc, doc_idx10,
ts, graph), and the lifecycle table builds with every declaration noted.
l5_lifecycle.main now routes SurrealDB before the generic _server branch,
which would have sent a served SurrealDB arm to ArcadeDB's served module.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* l2 (re-pin prep): ArangoDB's write checks its anchor first, and its update and delete are no-ops when the person is absent

The graph and cross-model fork found MongoDB writing a person whose
anchor the capped LDBC slice had not loaded, where every Cypher
engine's MATCH creates nothing; ArangoDB's write had the same
unconditional INSERTs. It now reads the anchor with DOCUMENT() and
FILTERs, still one AQL query. That made the update and delete reach an
absent person, where `UPDATE {_key}` and `REMOVE {_key}` raise "document
not found" and Cypher does nothing, so both now FILTER on _key (the
primary index) instead. Smoked at micro and on the capped SF1 slice:
all eight digests equal LadybugDB's on both.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* re-pin prep: Neo4j's vector quantization set and read back, a SCALAR arm, Memgraph's delete disclosed, SurrealDB served off lifecycle (DECISIONS #135)

Neo4j 2026.08.1 builds a vector index BINARY-quantized when the
definition names none, and both Neo4j arms named none while recording
fp32 (BUGS F164). Both definitions now set `vector.quantization.type`
(NONE on the fp32 arms, SCALAR on the new neo4j_dense_int8), read the
applied index configuration back onto the row (neo4j_vector_quantization,
neo4j_vector_index_config), and refuse the cell on a mismatch. The e2
lane now merges row_extra after build() as well as at connect, since the
read-back exists only once the index does. Smoked on the laptop: dense
micro NONE recall 0.9986 and SCALAR 0.991, e2 hybrid NONE.

The dense table says under it that Memgraph's timed delete includes the
FREE MEMORY it needs (its vector index keeps deleted points until GC);
COMPARATORS lists SurrealDB served under "Not added, and why" (its server
has no open or close of a database to time) and records the Neo4j
quantization history.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* re-pin prep: SurrealDB served runs sync=never and sync=every read back (F165); LadybugDB 0.21.2, pinned harness libraries, FalkorDB graph arm at 6.0.1 (#129)

SurrealDB 3.2.4 served takes its sync mode on the storage path; the October
campaign ran its default, a sync at every commit, in both durability classes
under a "not verified" label (BUGS F165, DECISIONS #136). The five served arms
start with ?sync=never, the strict class rewrites it to ?sync=every, the runner
reads "Sync mode:" back from the startup log before the client starts and
refuses the cell on a mismatch (surreal_sync_mode), bench_common moves the arm
from NO_DURABILITY_SETTING to STRICT_OF, and UNVERIFIED_ALLOWED is empty.
Laptop smoke through the runner, TPC OLTP micro: relaxed reads back never,
strict every transaction commit, reads unchanged between the two.

build_images.sh pins numpy, pandas, pyarrow, requests, and psycopg (bare in
every image; dbbench:duckdb had drifted to pandas 3.0.5) and moves ladybug to
0.21.2 (0.20.4 crashes the dense mutation phase). The FalkorDB graph arm moves
to the dense arm's 6.0.1 digest. Smoke at micro: graph OLTP in both classes,
graph analytics, and LadybugDB dense with the mutation phase, 10 of 10 ok,
13 answer groups agreeing with Neo4j. CAMPAIGN section 7 rows 47 and 48.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* l3_sparse: Qdrant's settle waits for every point it was sent, not only for green

The Qdrant sparse arm sends its batches with wait=False and only the last with wait=True, then waited for a green status. On the laptop neither meant the collection held the corpus: at the Big-ANN 100k slice the exact count at the first green was 66,064 of 100,000 (81,064 on the uint8 arm), and the lane searched it, recall@10 0.93 against exact ground truth (brute force over the same files agrees with the ground-truth file, 1.0). The same ingest with wait=True on every batch returns 1.0.

post_build now polls the exact count until it equals what was sent, with green, inside the build timer, bounded by BENCH_QDRANT_SETTLE_S (3600 s; a short collection refuses the cell). The row records qdrant_points_at_first_green and qdrant_settle_after_green_s. After the fix the arm returns 1.0 at that slice. The sparse lane and its multipass driver now merge an adapter's row_extra onto the row, as the dense lane does.

The mini rows at the October pin record 1.0 at 100k and 1M and 0.9998 to 0.9999 at 8.84M, so the race cost no recall there that a row shows; whether their build_s stopped before the last points were applied is not on any row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* re-pin prep: the quantization survey's int8 counterparts, each set and read back (DECISIONS #135)

Four dense arms and one sparse arm, each at its sibling's operating point, each setting its representation in the index definition, reading the applied one back from the engine onto the row, and refusing the cell on a mismatch (BUGS F164):

- mongodb_dense_int8: the vectorSearch field's quantization "scalar"; mongodb_dense keeps "none", now read back too (mongot_vector_quantization).
- memgraph_dense_int8: scalar_kind "i8"; memgraph_dense "f32", both read from SHOW VECTOR INDEX INFO (memgraph_vector_scalar_kind).
- lancedb_dense_fp32: IVF_HNSW_FLAT, unquantized. lancedb_dense keeps its name and its IVF_HNSW_SQ (int8), and the stale premise that SQ is LanceDB's only HNSW offering leaves the adapter, runner, and paper-table comments. Both read the index type and details back (lancedb_index_type, lancedb_index_details).
- arangodb_dense_int8: the FAISS factory "IVF<nLists>,SQ8" at the fp32 arm's nLists, nProbe calibrated to the same target. The pinned 3.12.11 accepts, trains, and answers from it (laptop probe: 20,000 SIFT vectors, nLists 566, recall@10 at nProbe 64 plain 0.9985, SQ8 0.9905); it validates factory strings only from 3.12.12, so the read-back (ivf_factory, a new IVF field) and the training state are the evidence.
- qdrant_sparse_uint8: SparseIndexParams datatype uint8; qdrant_sparse now names float32 and reads it back (qdrant_sparse_datatype).

Registered on their lanes (runner BACKENDS, LANES, MP_LABELS), in export_web's display, precision, dense overlay and sparse multipass lists and notes (the ArangoDB note and Memgraph's delete note match either arm), claims_check aliases, comparator_pins_check, make_paper_tables labels, and oct_repin_smoke.sh. degree_stamp strips an _fp32 suffix as it strips _int8. The l3d DURABILITY map gains the three int8 arms whose sibling has an entry, neo4j_dense_int8 among them: it was missing from its own commit and would have recorded the ingest-only note.

Laptop smoke through the runner, every cell ok, recall@10 fp32 / int8 at 5,000 and 20,000 SIFT vectors: MongoDB 0.9996 / 0.969 and 0.9977 / 0.9699; Memgraph 0.9997 / 0.9423 and 0.9992 / 0.9361; LanceDB 0.9999 / 0.9902 and 0.9994 / 0.9871; ArangoDB, each calibrated to the fallback 0.95, 0.9402 / 0.9465 and 0.9486 / 0.9447. Qdrant sparse fp32 / uint8: 1.0 / 1.0 on the synthetic 5,000, 1.0 / 0.9952 on the Big-ANN 100k slice. F7 and F14c pass on the smoke rows. QUANTIZATION.md and COMPARATORS.md record what each arm builds and reads back.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* l5 (re-pin prep): LadybugDB, Chroma, LanceDB, and sqlite-vec embedded on the lifecycle table (DECISIONS #131 item 5, row 40)

l5_lifecycle_embedded.py runs the lane's situations, sizes, session, and
mode set on the four other in-process engines on the page, each in the
image its other arms use. LadybugDB builds empty, doc, graph, vector (the
vector extension's HNSW, cosine, ml 32, mu 16), and ts, its threads and
buffer pool fitted as on its other arms (F160) and read back; Chroma
builds empty and vector; LanceDB builds empty, doc, doc_idx10 (BTree),
ts, and vector (IVF_HNSW_FLAT, unquantized); sqlite-vec builds empty and
vector (vec0, an exact scan, no drop cycle). Every other situation is
declared with the engine's own answer: LadybugDB's binder refuses a
secondary index and its projected graph is gone after a reopen, Chroma
stores no record without an embedding and has no local sparse index,
LanceDB's local connection runs no query language, and vec0 needs a
vector column and cannot be indexed.

An open of Chroma or LanceDB takes a handle on every collection or table
the database holds; LanceDB exposes no close, so its session drops its
handles. Durability is read back where the engine answers and named from
strace where it cannot: Chroma syncs at every add (8 fsync each) and
LanceDB never does (0 over 250 commits), neither with a setting, so
bench_common gains both strings and fairness_check names them, with
duckdb_lifecycle, which the gate would also have refused. LadybugDB
re-straced at 0.21.2: one fdatasync per commit, as at 0.20.4.

Laptop smoke through the runner, lc10k and lc100k, 64 cells clean; the
doc, doc_idx10, graph, and ts digests equal ArcadeDB's at both sizes.
The page prints one sentence naming the engines with no JVM start and one
per engine for what its session is (export_web OCT_PROSE); COMPARATORS,
FAIRNESS F6 and F10, and PAGE-SPEC say the same.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* export_web: the lifecycle table's declared situations fold to one sentence per engine

Each engine's situations it cannot build print as one sentence, grouped
by reason, each reason the head of the adapter's own declaration (up to
its first colon, without the statement tried, the engine's answer, or a
parenthetical such as a documentation link or a filing number). One
sentence per situation printed 21 near-identical sentences for the four
new engines alone. The statement and the answer stay on the row and in
the CSV; one lead sentence says so. The coverage gate gets one whole-row
absence per engine, as before.

The sqlite-vec and SQLite declarations gain a short head clause where
theirs ran to a full sentence, so the head reads as a reason: SQLite's
analytical view now groups with its graph under "no graph query
language", and its sparse reason drops "in any extension this arm loads"
from the head.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* l3d: Chroma's close note was wrong; chromadb 1.5.9 has Client.close(), and nothing the lane reads depends on it

The comment and the row's close_note said chromadb exposes no close,
from dir(chromadb.Client), which is the factory function. Client.close()
exists and stops the client's system. The arm still does not call it:
the lane never reopens, close_s is not on the page, and the client disk
reading is taken after the process exits, when the footprint is the same
with and without close() (laptop, 5,000 vectors, and 20,000 with 300
inserts and 300 deletes; every add() is persisted when it returns). The
note now says close_s is 0.0 by omission, not by measurement.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* re-pin prep: Chroma's dense arm calls Client.close(); LanceDB's no-close note names close_lsm_writers and why it does not apply

Base.close's rule is that close_s 0.0 means nothing to release, never "we did
not ask": chromadb 1.5.9 has Client.close(), so the Chroma dense arm now calls
it (0.002 s at micro) instead of recording 0.0 by omission. lancedb 0.39.0's
connection has no close, but a table has close_lsm_writers(), which drains
MemWAL writers that only merge_insert opens under an LSM write spec (read in
the pinned dbbench:dense image); the lifecycle arm appends with add(), so it
would be a no-op, and the note, docstring, page sentence, and COMPARATORS now
say exactly that rather than "neither exposes one". The lifecycle withheld
sentence gets its final period. Integration smoke on this tree: chroma_dense
and mongodb_dense_int8 at micro with the mutation phase, lancedb_lifecycle at
lc10k, 10 of 10 ok.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* l1_tpc: the served documents arm's analytics and point read go through POST /query (BUGS F163, CAMPAIGN row 25)

POST /command wraps even a SELECT in an auto-commit transaction, inside which
a scan stays on one thread, while the embedded twin reads with none. Since
ArcadeData/arcadedb#8792 (7f2c770697) POST /query runs an idempotent
statement without one. EXPLAIN on 26.10.1-SNAPSHOT build 6c2a72417a, the same
filtered GROUP BY: POST /query plans FETCH ... WITH FILTER (parallel) and
CALCULATE AGGREGATE PROJECTIONS (parallel), POST /command both sequential,
same answer (.notes repros/query-aggregate-scan/explain-6c2a72417a.txt).
Writes and the untimed verification scans stay on /command. Laptop smoke on
that snapshot with a 26.10.1.dev0 wheel: l1tpc olap and oltp at micro, served
and embedded, 11 answer groups agreeing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* re-pin prep: Neo4j's search expansion set to 1.0 and read back; ArangoDB 3.12.12 and Neo4j 2026.09.0 checked for the re-pin

Each Neo4j quantization type brings its own vector.default_search_expansion_factor
(NONE 1.0, SCALAR 1.5, BINARY 3.0 on 2026.08.1 and 2026.09.0), the candidate
multiple searched before re-scoring. The lane matches search effort at ef 100,
and Neo4j, with no per-query ef, asks for 100 candidates and keeps 10, so the
default would give the SCALAR arm 150 where its fp32 twin and every other engine
search 100. Both dense arms and the cross-model arm set 1.0 and refuse a
read-back that differs (neo4j_vector_search_expansion), as the Elasticsearch
int8 arm sets the minimum oversample. Laptop smoke: neo4j_dense 0.9986 and
neo4j_dense_int8 0.9908 recall at micro, neo4j_e2 ok, all reading back 1.0.

QUANTIZATION.md: Neo4j 2026.09.0 keeps the BINARY default; ArangoDB 3.12.12
accepts IVF<n>,SQ8 as 3.12.11 does and refuses an invalid factory at creation
(3.12.11 built an unusable index that answered by an exact scan). page_check
declares the new index-parameter fields (neo4j_vector_search_expansion,
ladybug_ml, ladybug_mu, lance_nprobes) with their family.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* export_web, page_check: three gaps the 26.10.1 publish rehearsal found

1. The skeleton's "no row on this table" sentence listed every registered
   backend without one, including OFF_PAGE_ARMS (postgres_tuned, the sparse
   no-settle ablation), which page_check's OFF-PAGE check then refused: a
   sentence naming an arm the page does not print. Off-page arms are excluded.
2. The sparse table relabelled its single pass as cold only when the second
   pass (l3smp) was there to fold in, so a skeleton (which declares l3smp
   absent) or a sparse landing ahead of its overlay printed "p50 ms" and
   coverage A1 refused it (MISS 'cold p50 ms', 12 undeclared cells). The lane's
   own pass is the first after the build either way; it is labelled cold, with
   no warm columns.
3. page_check's no-ArcadeDB-row-lost check counted a table the skeleton
   declares absent (skeleton_absent_tables: e4) as LOST once the live preview
   carried it. A declared skeleton absence is now named and skipped; a campaign
   payload carries no such list, so a real loss still fails.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* export_web, page_check: the rehearsal's coverage and condition findings on the 26.10.1 arms

Rehearsing the 26.10.1 publish on a skeleton freeze of 233 laptop rows, the
16 arms new since the October pin included (int8 counterparts, lifecycle
engines, cross-model arms, plain PostgreSQL time series, PG+AGE graph):

- coverage A2: 19 numeric fields the new arms record were neither a column nor
  declared and would have refused the real publish. Declared with reasons:
  PostgreSQL's SHOW read-back (pg_*), Elasticsearch's processors and heap,
  duckdb_threads, lc_affinity_cpus, vector_index_size (a build check), the
  Qdrant settle pair (BUGS F166), setup_s (the residue the dense split note
  states as a range, BUGS F101), the ArcadeDB documents loader configuration
  (async_writers, lineitem_buckets, load_call_rows, served_load_batch), and
  es_num_candidates and es_rescore_oversample with the index parameters.
- conditions: the answer-check note quoted each declared absence's raw engine
  error, whose digits (SurrealDB's parse position [1:8]) have no source, and
  with every lifecycle engine declared it ran to 11,000 characters. It now
  keeps each reason, drops the attempted statement and the engine's answer
  (they stay on the row), and names every engine sharing a reason once; the
  lifecycle note is 5,273 characters. vec0 joins Neo4j as an exempt name.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* equivalence_check, export_web: the lifecycle read compared by deployment, and its note says only what it compared

The lifecycle read was declared NOT_COMPARABLE because the served arm runs a
different mode set, so its post-state counts differ by construction. With six
embedded engines on the table that declaration exempted all of them: an
embedded engine answering differently from ArcadeDB embedded would have been
skipped, not refused. equivalence_check now splits the lifecycle group by
deployment (SPLIT_BY_DEPLOYMENT; E3 counts silence only within a deployment),
so embedded arms are held to ArcadeDB embedded's answer and the served arm to
its own deployment. Rehearsal (233 laptop rows): documents, documents with ten
indexes, the graph, and time series agree across 4 to 6 engines each; nothing
else in the gate's output changed.

export_web: the lifecycle table's answer-check note listed every declared
absence with its reason, restating the per-engine absence sentences, 5,273
characters. It is now generated from the gate's groups and says only which
situations were compared and how, which had no second answer, which were not
compared and why, and why the server rows stand apart: 756 characters, no
engine error text. Every other table's note is byte-identical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* make_2610_stages: the 26.10.1 measurement's queue scripts, in DECISIONS #133's order

A profile of make_october_stages, not a copy: its template and emit() are
reused, and what differs is replaced by asserted exact-string substitution.
Eleven stages: graph interactive (both classes), graph analytics, cross-model
(both classes), documents, dense with the multipass overlay, sparse with the
second pass, time series, the e4 decomposition, lifecycle for ArcadeDB and
SurrealDB, the Python-cost table (host-side, run_bench.sh's six steps, five
runs, output outside the tree), and last the lifecycle expansion (row 40). The
first waits on qOA5 ALL-DONE.

The pin is a release pair named at generation (ARCADEDB_WHEEL by file name and
sha256, ARCADEDB_SERVER_IMAGE, ARCADEDB_ENGINE_COMMIT); the generator refuses to
emit without all three, or for a pre-release wheel unless --allow-prerelease
(DECISIONS #42). Each stage carries its comparators' digests and refuses to
start if runner.py has moved one. cell_cost_check aborts only on this pin's own
rows; an unmeasured tier is scored from October's rows and only warns (at tpch10
that puts MongoDB at 103% and SurrealDB served at 106% of a cap October met).

check_coverage() derives from runner.LANES and export_web._TABLE_LANE: every
registered arm of every page lane in exactly one stage per workload (passes;
the l1 tabular lane feeds no page table). --project sums October's measured
cells per stage: 310 h measured + 243 h estimated = 553 h.

verify_pair_c25.sh takes PAIR_IMAGE; run_bench.sh takes ONLY_STEPS,
PYCOST_RESULTS, PYCOST_DBS, and UV_RUN_ARGS, all opt-in. Dry emission with the
snapshot pair: bash -n clean, queue_lint 11 scripts 0 problems.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr

* re-pin prep: the lifecycle lane's roster is the embeddable multi-model engines; the expansion stage is dropped (DECISIONS #139)

The lifecycle table compares the session cost of ArcadeDB and SurrealDB, the
two multi-model engines a process can open and close in-process. #131 item 5
had extended it to the single-model embedded engines, which contradicts #131
item 1 (a single-model specialist runs only on its own model's tables), and
the user took the recommendation to keep it multi-model. runner.LANES drops
sqlite, duckdb, ladybug, chroma, lancedb, and sqlite_vec lifecycle from the
roster (their BACKENDS entries and adapters stay, runnable by name);
make_2610_stages loses qRK; COMPARATORS lists them under "Not added, and why";
PAGE-SPEC's lifecycle row says so. Coverage check: 0 problems, lifecycle
roster 3.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Lj5isEaoaGCeaZe5nNoWX4

* l6_restart: the server-restart lane, every served engine restarted in place on one model's data (DECISIONS #139 item 2)

A served engine has no open or close of its own, so the lifecycle table cannot
measure it; it has a restart instead. l6_restart.py stops and starts the cell's
OWN server container through the Docker socket (docker_api.py, stdlib only;
the runner mounts the socket for this lane alone), so data, flags, cpuset,
cap, and network name never change. Per cycle: a clean stop after a session
that wrote nothing (the image's STOPSIGNAL, then SIGKILL after a grace no
engine reaches, recorded if it ever fires), a restart timed to the first
liveness call (start) and to a fixed read answering exactly as before the stop
(first answer), a committed write batch, a stop after it, a restart after it,
and the batch read back or the cell fails. Fresh clients per probe, so no
driver backoff is timed; warm OS page cache, stated.

16 arms, one model each through its own lane's loader at that lane's tiers and
envelope: documents (ArcadeDB, SurrealDB, ArangoDB, MongoDB, PostgreSQL),
graph (ArcadeDB, Neo4j, Memgraph, FalkorDB), dense (ArcadeDB, Qdrant, Milvus,
Elasticsearch, MongoDB + mongot), time series (ArcadeDB native, QuestDB).
ArcadeDB's server runs all four because a page tier without ArcadeDB is
withheld. PostgreSQL stands for pgvector, PG+AGE, TimescaleDB (one binary).

Registered: runner LANES + socket mount + env allowlist; export_web table,
scale labels, _TABLE_LANE, SOURCES, registered sentences (its own durability
sentence instead of the relaxed-match note); make_paper_tables tiers;
page_check NOT_PRINTED; fairness_check producer; equivalence_check splits the
lane by model (the skeleton runs every model at micro) and declares the vector
read not comparable across engines; make_2610_stages qRK-qRN, coverage 0
problems, queue_lint clean; PAGE-SPEC, PROTOCOL, COMPARATORS, FAIRNESS F10c.
dbbench:mongo-search's entrypoint now forwards the stop signal to mongod and
mongot (bash as PID 1 ignored it).

Laptop smoke, all 16 at the smallest tier: every cell ok, every batch read
back, answers agree within each model; the shutdown trace ran on all 16.
Publish rehearsal (skeleton): the table builds; the gates refuse only for the
laptop host line and the mixed laptop pins (ArcadeDB 802233cb57 smoke rows,
FalkorDB 6.0.1 beside October's 4.20.6).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Lj5isEaoaGCeaZe5nNoWX4

* harness: every ArcadeDB server version read uses GET /api/v1/server?mode=basic (ArcadeData/arcadedb#8909)

The default mode exports rate metrics, and MetricMeter's first ask loops once
per second since 1970 (about 2 s per fresh meter), so the first call after a
server start takes about 20 s and the second about 2 s on every build from
26.8.1 to 802233cb57 (filed upstream as #8909, Java repro in .notes). Every
lane read the version with it in connect(): outside every timer, so no
October number moved (all 305 served ArcadeDB rows at the pin recorded their
version), but the restart lane's fresh client did it before its first query,
and its ArcadeDB dense cell read 22 s. mode=basic returns the version in
33 ms. Re-smoke of that cell: first answer 1.04-1.19 s over five cycles.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Lj5isEaoaGCeaZe5nNoWX4

* make_2610_stages: the restart table's dense stage runs 1M and 10M (DECISIONS #139 item 2)

A vector index's restart cost is the index coming back, and it shows at 10M;
the lane already supports deep10m (l6_restart DENSE_DATA), so qRM runs small
and deep10m and the freeze admits the restart lane's deep10m rows. About
25-40 h more, mostly loads, on top of the table's 17-25 h. Repetitions stay
at five (#133). Also fixes a comment that still read "No qRK" after qRK
became the documents restart stage. Coverage 0 problems, queue_lint clean on
the 14 generated scripts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Lj5isEaoaGCeaZe5nNoWX4

* export_web: the restart table names an engine whose clean stop ends in an abort, and labels the 10M dense tier

The restart lane records each stop's exit code. 0, or 143 for a process that
ends by the stop signal after its shutdown hook (every JVM here, and Qdrant),
is a normal end; Milvus ends every clean stop in a panic in its shutdown path,
exit 134, with its data intact (filed as milvus-io/milvus#53947, reproduced on
3.0.1 and 3.0.2 with the stock standalone config). The table now says so,
generated from the rows: on the 34 restart smoke rows only Milvus is named.
Also adds the (restart, deep10m) scale label, "dense vectors, 9.99M (Deep)",
which the 10M tier of f62f7e17d3 needed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Lj5isEaoaGCeaZe5nNoWX4

* CAMPAIGN section 7: rows 37, 38, 39, and 41 are built and smoked on repin-prep (de82a23762, 5cc6f56cf4, 4eb14e4009), not 'decided, not built'

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL

* l4: hold the parsed TSBS corpus at 89 bytes a point and release it after ingest (CAMPAIGN 7 row 18, BUGS F66)

WHAT HELD THE POINTS. Nothing downstream of a chunk: the lane never read
the corpus in chunks. l4_tsbs.parse_lp built the whole corpus as one list
before the ingest timer (deliberately, so the parse stays out of every
arm's ingest_s), main() kept `pts` referenced until the cell exited, and
TS_CHUNK and the arms' batch sizes only slice that list. Every point owned
a fresh 5-tuple plus a fresh host string, a fresh timestamp int, and three
fresh floats: 281 bytes a point, linear in the points parsed (tracemalloc
at 500k points: l4_tsbs.py:152 tuple + int 59.3 MiB, :150 host 22.8 MiB,
:153-155 one float each 11.4 MiB). DuckDB's arm also kept its registered
Arrow copy of the corpus alive through every query.

THE FIX keeps the parse outside the timer and the corpus whole during the
ingest. Streaming the parse into the ingest would either charge it to the
timer or, with the timer paused around it, give the arms whose engine works
after the call returns (ArcadeDB's native async executor, the COPY arms,
QuestDB's ILP socket) uncharged background time.
- parse_lp shares the repeated values (100 or 1,000 hosts, 25,920
  instants, readings 0-100) by their text, so a point costs only its tuple
  and list slot. Same values, types, order, and skipped lines, checked
  point for point against the old parse on ts100 (2,592,000 points),
  ts1000 (25,920,000), and a file of malformed lines.
- main() drops the corpus after the ingest timer and after the engine's
  own settle, and DuckDB unregisters its Arrow table there
  (release_ingest_input). The lane's settle floor absorbs the release, so
  the gap before the first query is unchanged; the row records
  corpus_release_s, which page_check's NOT_PRINTED now classifies.
No timer covers different work, no arm sends anything different, and no
batch size changes.

MEASURED ON THE LAPTOP (memory only; no timing claims):
- RssAnon held by the parsed corpus, client image, Python 3.12:
  ts100 695 -> 221 MiB, ts1000 7,305 -> 2,186 MiB (281/296 -> 89 B/pt).
- runner smoke, ts100, BENCH_QITER=5, client_peak_anon_mib before -> after:
  sqlite 722.5 -> 237.4, duckdb 1018.9 -> 543.3,
  arcadedb_ts_doc 9347.8 -> 8854.2, arcadedb_ts_native 3846.5 -> 3729.2.
- sqlite driver sampled every 0.5 s: before, it climbed to 706 MiB during
  the parse and stayed there through ingest and queries; after, 233 MiB
  through ingest and 52 MiB through the queries.
- surrealdb_ts embedded at ts100 (sampler, QITER=1): RssAnon peak during
  ingest 3,580.6 -> 3,105.1 MiB, overall 4,252.2 -> 3,651.9 MiB.
- Answer digests: all six queries equal before and after on sqlite,
  duckdb, arcadedb_ts_doc, arcadedb_ts_native, and surrealdb_ts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL

* CAMPAIGN section 7 row 18 and FAIRNESS F1: the TS driver never chunked its parse; the fix (5f61f963fe) is built, row 18 says what was found and measured

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL

* Disclose Elasticsearch's two overrides and add the registry that holds every override disclosure (CAMPAIGN 7 row 21)

PROTOCOL section 7 said NOWHERE for xpack.security.enabled=false and number_of_replicas=0: no row field, no sentence under any table. Both are now asked of the engine and stamped on the row, es_security_enabled from `_xpack` and es_replicas from the index settings, by the sparse adapter and by the dense adapters. The dense read goes through a new readbacks() hook that main(), the dense multipass driver, and the restart lane call after the build timer has stopped, so no read-back sits inside a published timer; the restart lane merges only the registered fields.

Every October table that shows an Elasticsearch arm now carries a generated sentence saying security is off and the index has no replica. The sentences come from overrides.py, a registry with one record per override (the arms that run it, the row field and the value the sentence claims, the sentence, and the regexes the printed sentence must satisfy); export_web appends them, and the last column of PROTOCOL.md section 7 cites `override: <key>`.

Checked three ways. page_check fails a table that shows an Elasticsearch arm and prints no sentence naming the override. fairness_check F15 fails a 2026-10 row of such an arm that lacks the field or carries another value. test_overrides.py reads PROTOCOL.md section 7 and fails when a registered key is cited by no row or a row cited as disclosed still says NOWHERE, and it builds a table without the sentence and a row without the stamp to show each gate fails. Every new numeric field is declared in NOT_PRINTED, which page_check's field coverage requires. Laptop smoke through the runner: l3s, l3d, restart, and multipass-overlay Elasticsearch records carry es_security_enabled false and es_replicas 0.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL

* Disclose DuckDB's thread pool and its experimental HNSW persistence (CAMPAIGN 7 row 21)

DuckDB runs with PRAGMA threads fitted to the cell's cpuset and, on its vector arms, with hnsw_enable_experimental_persistence switched on. Neither was named under any table. Each is now read back from the engine with current_setting and stamped: duckdb_threads on the documents, time-series, dense, and cross-model arms (the graph arm's duckpgq_threads and the cross-model arm's duckdb_threads used to record the number passed and now record the engine's answer) and duckdb_hnsw_persistence on the dense and cross-model vector arms. Every read happens at connect, outside every timer. The lifecycle DuckDB arm is not on a page table (the single-model embedded engines are off its roster), so it carries no stamp.

Two generated sentences appear under every table that shows DuckDB. fairness_check F15 holds the thread count to the size of the row's own cpuset, so a PRAGMA that did not take fails on the row: that is the claim "fitted to the cpuset" made checkable. page_check fails a table that shows the arm without the sentence, and test_overrides.py proves both gates fail on a 20-thread row in a 12-CPU cell and on a table built without the sentence. Laptop smoke through the runner (cpuset 4-15): l1tpc, l4, l3d, and e2 DuckDB rows read 12 threads, and the vector arms read the persistence flag true.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL

* Disclose Neo4j's page cache and checkpoint interval (CAMPAIGN 7 row 21)

The runner fits Neo4j's page cache to the cell (its image would set one small fixed size) and sets the checkpoint interval to 5 seconds on the graph arm and the composed cross-model stack. Neither was named under any table. Both are now read back from the engine with SHOW SETTINGS at connect: neo4j_pagecache on the graph, dense, and cross-model arms and the restart lane's graph model, and neo4j_checkpoint_interval with the engine's own default beside it on the two arms that set it. The registry lists exactly those two for the checkpoint, so a table that shows only the dense or the single-engine cross-model arm does not claim it.

The sentences come from the rows: the checkpoint sentence quotes the engine's own interval and default, and falls back to wording with no digits if the rows lack them. fairness_check F15 holds the engine's page cache to the size the cell passed (`server_pagecache`, from docker), so an arm that ran the image's fixed default fails. page_check fails a table that shows a Neo4j arm without the sentences. Laptop smoke through the runner: l2, l3d, e2, and restart Neo4j rows read a page cache equal to the cell's (3.00GiB against 3.0g at micro), and the two arms that set the checkpoint read 5s against a 15m default.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL

* Disclose ArcadeDB's addHierarchy on the dense vector index (CAMPAIGN 7 row 21)

The dense lane builds ArcadeDB's LSM_VECTOR index with addHierarchy true where the engine default is a flat graph. The embedded arm now reads the setting back from the engine (the per-bucket index serialises it) and stamps arcadedb_add_hierarchy with the source "index metadata read back from the engine". The served arm cannot be asked: its schema queries return no index metadata at this pin, so it records the CREATE INDEX statement it sent, parsed from the exact text, with the source "requested in the CREATE INDEX statement", so a reader can tell a read-back from a request. Both reads go through the readbacks() hook, after the build timer.

A generated sentence under every table that shows a dense ArcadeDB arm says the index is a layered graph, the structure the hnswlib family uses, and that the default is flat; it claims no effect on a number, which has not been measured. fairness_check F15 requires the field and the source on every dense ArcadeDB row, and page_check fails a table without the sentence. Laptop smoke: embedded rows read true from the engine, served rows true as a request, in the lane rows, the multipass overlay records, and the restart row.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeKEhHXJTy1H1GgbfBZ8mL

* Disclose the served ArcadeDB's query-memory limit (CAMPAIGN 7 row 21)

Every served ArcadeDB arm is launched with arcadedb.queryMaxHeapElementsAllowedPerOp fixed at 5,000,000; the embedded package leaves the engine default, which scales with the heap, so the two deployments can run different limits, and the deployment table compares exactly those. The runner now asks the engine for the value it runs with (`SELECT FROM schema:database` over HTTP, from the host, before the client starts and outside every timer) and stamps server_query_max_heap_elements on every served ArcadeDB row, whichever lane, with server_query_max_heap_source naming how it was read; if the host cannot reach the container the container's own JAVA_OPTS value is recorded and the source says so. One runner function covers every lane that has a served ArcadeDB arm, whose adapters have no common stamping path. The dense multipass overlay records inherit the field from their campaign row, as they do the other server fields.

A generated sentence quotes the stamped value under every table that shows a served ArcadeDB arm, and the artifact-backed deployment table takes the number from the runner's own launch configuration. The carriers are a stati…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

to_columns, to_dataframe, and to_arrow drop every property that the first row lacks, and export_to_csv raises after writing part of the file

1 participant