Search building blocks published under @statewalker/*: a pluggable indexer contract, full-text and vector indexers (in memory, DuckDB, PGlite), Markdown chunking, and content extractors that turn PDF, DOCX, XLSX, HTML and Markdown files into text. Applications pick one backend, create named indexes, put documents into named sub-indexes, and search them together.
indexer-api (kernel types + helpers)
/ \
indexer-fulltext indexer-vector (modality types, providers, access handles)
\ /
indexer-core (composite index, RRF, indexer builders)
/ | \ \
indexer-mem indexer-mem-flexsearch indexer-duckdb indexer-pglite
(vectors) indexer-mem-minisearch (SQL backends)
Side packages: indexer-chunker, content-extractors (no indexer dependency)
Private: indexer-search (SearchPipeline), indexer-tests (conformance suites)
| Package | What it gives you | npm |
|---|---|---|
| @statewalker/indexer-api | Indexer, Index, SearchIndex, ModalityProvider types; sub-query/sub-result helpers. |
npm |
| @statewalker/indexer-fulltext | Full-text types, FullTextProvider, newFullTextAccess, config helpers. |
npm |
| @statewalker/indexer-vector | Vector types, VectorProvider, newVectorAccess, config helpers. |
npm |
| @statewalker/indexer-core | Code for backend authors: composite index, RRF, persistence-backed and SQL indexer builders. | npm |
| @statewalker/indexer-mem | In-memory vector sub-index (MemVectorIndex) and memVectorProvider. |
npm |
| @statewalker/indexer-mem-flexsearch | In-memory indexer: FlexSearch full-text + in-memory vectors, optional persistence. | npm |
| @statewalker/indexer-mem-minisearch | In-memory indexer: MiniSearch full-text + in-memory vectors, optional persistence. | npm |
| @statewalker/indexer-duckdb | DuckDB indexer: BM25 full-text (fts) + HNSW vectors (vss). |
npm |
| @statewalker/indexer-pglite | PGlite indexer: tsvector/GIN full-text + pgvector HNSW. |
npm |
| @statewalker/indexer-chunker | Markdown-aware chunking before indexing. | npm |
| @statewalker/content-extractors | Text extraction from PDF, DOCX, XLSX, HTML, Markdown, plain text. | npm |
| @statewalker/indexer-search | SearchPipeline (expand, embed, search, rerank, cite), weightedBlend, query parser. |
private |
| @statewalker/indexer-tests | Vitest conformance suites that every backend runs. | private |
External @statewalker dependencies: @statewalker/db-api (runtime, indexer-duckdb) and @statewalker/db-duckdb-node (tests of indexer-duckdb).
- Install Node.js 24 and enable corepack:
corepack enable. pnpm 10 is pinned inpackageManager. - Install:
pnpm install. - Build all packages:
pnpm build. Tests and typechecks of a package use the builtdist/of its dependencies, so build first. - Test:
pnpm test. Type-check:pnpm typecheck. - Before pushing:
pnpm lint:checkandpnpm format:check(orpnpm lint/pnpm formatto fix).
For one package: pnpm --filter @statewalker/<name> <script>.
- Sub-indexes are named, not keyed by modality. One index can hold two full-text sub-indexes (English and French) or two vector sub-indexes (content and summaries). Queries and results are addressed by sub-index name, so the kernel never needs to know modality types; a new modality is a new package.
Index.searchfuses with parameter-free RRF. BM25 and cosine scores are not comparable, so the top-levelscoreis a rank-fusion score. The native score of each sub-index stays onresult.subResults[name]. Weighted blending is a caller-side step (weightedBlendinindexer-search).- Persistence is per sub-index. A sub-index that can save itself implements
serialise()/loadFrom(). The in-memory backends save through anIndexerPersistenceport; the SQL backends store everything in their database. - The SQL backends share one indexer builder (
createSqlIndexerinindexer-core) because they share a per-index docs table that maps paths to row ids; the in-memory backends use the provider-basedcreatePersistenceBackedIndexer. - Packages ship
dist/andsrc/.exportspoint atdist/; the sources are included for reading and debugging.
SearchRequest.pathsdoes not filter. The composite passes only each sub-query to its sub-index; the top-levelpathsis not forwarded. Putpathson each sub-query ({ queries, paths }/{ embeddings, paths }), or you get hits from the whole index.- Top-level
topKonly truncates the fused list. Each sub-index retrievessub-query.topKhits (default 100 in all current backends), then RRF fuses them. - DuckDB downloads extensions.
createDuckDbIndexerrunsINSTALL fts; INSTALL vss;. Offline or sandboxed hosts fail at init with DuckDB's extension download error. - PGlite rejects unknown languages with
pglite FTS: unsupported language "<x>" (must be an ISO-639-1 code or a Postgres text-search config name). languagedoes nothing in the in-memory full-text indexes. FlexSearch and MiniSearch run without stemming; the value is only stored.
| Command | Does |
|---|---|
pnpm build |
pnpm -r run build (tsdown, writes dist/) |
pnpm test |
Vitest in every package under packages/ |
pnpm typecheck |
tsc --noEmit in every package |
pnpm lint / pnpm lint:check |
Biome check with / without writes |
pnpm format / pnpm format:check |
Biome format with / without writes |
pnpm changeset |
Add a changeset (bump level and changelog text) |
Packages are published to npm from CI with changesets. After CI passes on main, changesets are generated for packages whose packed contents differ from npm, and a "chore: version packages" pull request is opened; merging it publishes with npm provenance. Run pnpm changeset in your pull request to choose the bump or the changelog text yourself. Dependency updates arrive as Renovate pull requests.
| Path | Contents |
|---|---|
packages/* |
Workspace packages |
pnpm-workspace.yaml |
Workspace globs and the dependency catalog |
biome.json |
Lint and format rules |
turbo.json |
Task graph |
CONTEXT.md |
Vocabulary of the indexer contract |
MIT. See LICENSE.