diff --git a/README.md b/README.md index 1c91f9b..e8db06e 100644 --- a/README.md +++ b/README.md @@ -1,20 +1,115 @@ -# statewalker-notebooks +# statewalker-knowledge -Browser-based Observable notebooks built on the statewalker stack. +## What it is -## Packages +A pnpm workspace of packages that turn notebooks into web pages that run in the browser. A +notebook is a Markdown file or an [Observable notebook-kit](https://github.com/observablehq/notebook-kit) +HTML document. `@statewalker/notebook-build` builds a tree of notebooks into a site of +executable pages, `@statewalker/notebook-db` backs SQL cells with any `@statewalker/db-api` +database, and `@statewalker/notebook-site` serves the built site from one fetch handler. All +storage goes through `FilesApi` (`@statewalker/webrun-files`), so the same code runs under +Node, in a Worker and behind a ServiceWorker. -| Package | Description | -| --- | --- | -| [@statewalker/notebook-build](packages/notebook-build) | Builds a tree of notebooks on a FilesApi into a static site of executable pages, incrementally. | -| [@statewalker/notebook-db](packages/notebook-db) | Backs notebook-kit SQL cells with any `@statewalker/db-api` database: a live client, and a build-time precompute that writes the JSON notebook-kit's cached client fetches. | -| [@statewalker/webrun-http-events](packages/webrun-http-events) | Generic publish/subscribe over Server-Sent Events: a standard FetchHandler and a matching client. | -| [@statewalker/notebook-site](packages/notebook-site) | Serves a built notebook site as a FetchHandler: pages and attachments from files, modules from a module server in hosted mode or from the export in static mode, plus the rebuild event stream. | +## The shape + +``` +packages/ + notebook-build/ notebooks (FilesApi) -> pages + attachments + module closure (FilesApi) + notebook-db/ db-api Db -> notebook-kit SQL source; build-time SQL cache files + notebook-site/ built site (FilesApi) -> one SiteHandler (Request -> Response) +apps/ + demo/ private: builds ./notebooks and serves them on localhost:8099 +``` + +How the pieces connect: -## Development +``` + notebooks/*.md, *.html ──► newNotebookBuild ──► output/ (pages, attachments, + (FilesApi) │ ▲ static: /_m/ closure) + │ └── moduleServer (@statewalker/webrun-modules) + ▼ + cache/ (serialized notebooks, incremental state) -```sh -pnpm install -pnpm run build -pnpm run test + Request ──► newNotebookSite ──► /_m/* moduleServer (hosted mode only) + ├─► /_events/* PubSub (SSE) (optional) + └─► /* output/ ``` + +| Package | What it does | Published | +| --- | --- | --- | +| [@statewalker/notebook-build](packages/notebook-build) | Builds a tree of notebooks on a `FilesApi` into a static site of executable pages, incrementally. | [npm](https://www.npmjs.com/package/@statewalker/notebook-build) | +| [@statewalker/notebook-db](packages/notebook-db) | Backs notebook-kit SQL cells with any `@statewalker/db-api` database: a live client, and a build-time precompute that writes the JSON files notebook-kit's cached client fetches. | [npm](https://www.npmjs.com/package/@statewalker/notebook-db) | +| [@statewalker/notebook-site](packages/notebook-site) | Serves a built notebook site as a `SiteHandler`: pages and attachments from files, modules from a module server (hosted mode) or from the export (static mode), plus the rebuild event stream. | [npm](https://www.npmjs.com/package/@statewalker/notebook-site) | +| [@statewalker/notebook-demo](apps/demo) | Builds the notebooks in `apps/demo/notebooks` and serves them. | private | + +## How to run it + +Requirements: Node.js 24 and pnpm 10 through corepack. The pnpm version is pinned by +`packageManager` in `package.json`. + +1. `corepack enable` +2. `pnpm install` +3. `pnpm run build` (each package builds with tsdown into `dist/`) +4. `pnpm run test` +5. Optional, browser tests: `pnpm --filter @statewalker/notebook-build exec playwright install chromium`, + then `pnpm run test:browser`. +6. Optional, the demo: `pnpm --filter @statewalker/notebook-demo start`, then open + . + +## Why it is the way it is + +- **Every store is a `FilesApi`.** Sources, output, build cache and module cache are all + `FilesApi` instances, so the build and the server run anywhere a `FilesApi` backend exists: + the Node filesystem, memory, OPFS. +- **Nothing reads a DOM global.** `notebook-build` takes a `document` and a `DOMParser` as + options (jsdom under Node, the natives in a browser), and `notebook-site` is compiled without + the DOM lib, so it can run in a Worker. +- **npm imports are served from the same origin.** In hosted mode a live module server + resolves and transforms npm packages on demand under `/_m/`; in static mode the build writes + the whole dependency closure into the output. In both cases a page makes no third-party + requests at run time. +- **Browser tests are a separate script.** `pnpm run test` never launches a browser. The browser + tests drive a real Chromium through Playwright from Node, because they have to execute built + pages, a real DuckDB-WASM and a real ServiceWorker. +- **Integrations are peer dependencies.** `@statewalker/webrun-files`, `webrun-builder`, + `webrun-modules`, `webrun-site-builder` and `@statewalker/db-api` are peers, so an + application and these packages share one copy of each interface. + +## What will surprise you + +- **The first browser test run or demo start is slow.** On a cold cache the module server + downloads and transforms the npm dependency graph of the notebooks (Observable Plot, DuckDB). + The browser tests allow up to 15 minutes for their setup hook for this reason. +- **`pnpm exec playwright` at the root is the wrong Playwright.** Playwright is a dependency + of the packages, not of the root; run it through `pnpm --filter exec` so the + browser matches the version the tests import. +- **The build keeps its scanner state inside the notebooks tree.** The build engine stores its + state in the `notebooks` `FilesApi` under `.notebook-build/`. The demo ignores + `notebooks/.notebook-build/` in git; do the same in your own source tree. +- **There is no `format` script.** `pnpm run lint` runs `biome check --write .`, which lints and + formats in one pass; `pnpm run lint:check` checks both without writing. + +## Reference + +### Commands + +| Command | What it does | +| --- | --- | +| `pnpm run build` | `tsdown` in every package | +| `pnpm run test` | builds each package, then `vitest run` | +| `pnpm run test:browser` | Playwright browser tests in every package | +| `pnpm run typecheck` | `tsc --noEmit` for sources and browser tests | +| `pnpm run lint` | `biome check --write .` | +| `pnpm run lint:check` | `biome check .` | +| `pnpm changeset` | add a changeset (bump type and changelog text) to a pull request | + +### Releases + +Packages are published to npm from CI with [changesets](https://github.com/changesets/changesets). +A pull request may carry a changeset made with `pnpm changeset`; otherwise one is generated for +each package whose packed contents differ from the version on npm. Each package ships `dist/` +(JavaScript and `.d.ts`) and its TypeScript sources in `src/`; `exports` point at `dist/`. + +### License + +MIT, see [LICENSE](LICENSE). diff --git a/apps/demo/README.md b/apps/demo/README.md new file mode 100644 index 0000000..2a86ef7 --- /dev/null +++ b/apps/demo/README.md @@ -0,0 +1,62 @@ +# @statewalker/notebook-demo + +## What it is + +A private app, not published. `serve.mjs` builds every notebook in `./notebooks` with +`@statewalker/notebook-build` and serves the result with `@statewalker/notebook-site` on a local +Node HTTP server. npm imports in the notebooks (`npm:@observablehq/plot`) are resolved and +transformed on demand by `@statewalker/webrun-modules` and served from the same origin under +`/_m/`, so the page makes no third-party requests at run time. + +## The shape + +``` +apps/demo/ + serve.mjs build once (or on change with --watch), then serve on $PORT (8099) + notebooks/index.md the demo notebook: reactive js cells, Observable Plot, a broken cell + .out/ built site (generated, git-ignored) + .cache/build/ build artifacts and incremental state (generated, git-ignored) + .cache/modules/ downloaded and transformed npm modules (generated, git-ignored) +``` + +## How to run it + +1. From the repository root: `pnpm install`. +2. `pnpm --filter @statewalker/notebook-demo start`. This builds `notebook-build` and + `notebook-site`, then runs `node serve.mjs`. +3. Open . Set `PORT` to use another port. + +To rebuild when a notebook changes, run `node serve.mjs --watch` from `apps/demo` after the +packages are built. + +## Why it is the way it is + +- **Hosted mode.** The demo passes `mode: "hosted"` and the module server to both the build and + the site, so modules are served live from `/_m/` instead of being copied into `.out/`. +- **jsdom for the DOM.** The build reads no DOM global; under Node the demo creates a jsdom + window and passes its `document` and `DOMParser`. +- **One deliberately broken cell.** `notebooks/index.md` contains a cell that does not compile, + to show that the error is displayed in place and the other cells still run. + +## What will surprise you + +- **The first start is slow.** It downloads and transforms Observable Plot and its dependency + graph into `.cache/modules/`. The console shows + `building (first run downloads and transforms the npm imports)…`. +- **`notebooks/.notebook-build/` appears in the source tree.** The build engine keeps its scanner + state in the `notebooks` store. It is git-ignored. +- **Build failures print `✗ undefined: undefined`.** `onFailed` receives an array of failures, + but `serve.mjs` destructures it as a single `{ notebookPath, error }`. Look at the page itself + for the error until this is fixed. + +## Reference + +| Command | What it does | +| --- | --- | +| `pnpm --filter @statewalker/notebook-demo start` | build the two packages, then `node serve.mjs` | +| `node serve.mjs` | build once and serve | +| `node serve.mjs --watch` | also rebuild when a file under `notebooks/` changes | + +| Environment variable | Default | Meaning | +| --- | --- | --- | +| `PORT` | `8099` | HTTP port | diff --git a/packages/notebook-build/README.md b/packages/notebook-build/README.md index 0469f22..028d125 100644 --- a/packages/notebook-build/README.md +++ b/packages/notebook-build/README.md @@ -1,150 +1,217 @@ # @statewalker/notebook-build -Turns a tree of notebooks on a [`FilesApi`](https://github.com/statewalker/webrun-files) into a -static site of executable pages, incrementally. +## What it is -A notebook is a Markdown file (fenced `js`/`ts`/`ojs`/`sql` blocks become cells, everything else -is prose) or an [Observable notebook-kit](https://github.com/observablehq/notebook-kit) HTML -document. Each one becomes an `.html` page that imports notebook-kit's runtime and runs its own -cell graph in the browser. +Turns a tree of notebooks on a `FilesApi` (`@statewalker/webrun-files`) into a static site of +executable pages, incrementally. A notebook is a Markdown file (fenced `js`/`ts`/`ojs`/`sql` +blocks become cells, everything else is prose) or an +[Observable notebook-kit](https://github.com/observablehq/notebook-kit) HTML document. Each one +becomes an `.html` page that imports notebook-kit's runtime and runs its own cell graph in the +browser. -```sh -npm install @statewalker/notebook-build -``` - -## Usage +## Why it exists -```ts -import { newNotebookBuild } from "@statewalker/notebook-build"; +notebook-kit ships a Vite plugin for building notebooks. This package builds them over +`FilesApi` instead, with no bundler and no DOM global, so a build can run under Node or in a +browser against any `FilesApi` backend. npm imports are resolved through a module +server and served from the site's own origin, so a built page makes no third-party requests. +The build is incremental: a run re-derives only the notebooks whose inputs changed. -const build = newNotebookBuild({ - notebooks, // FilesApi: the sources - output, // FilesApi: the site - cache, // FilesApi: the build's own state - moduleServer, // @statewalker/webrun-modules - dom: { document, parser }, // jsdom under Node; the natives in a browser - mode: "static", - stylesUrl: "/_m/@observablehq/notebook-kit@2.6.4/dist/src/styles/index.css", - onRebuilt: (changed) => deploy(changed), - onFailed: (failures) => failures.forEach((f) => console.error(f.notebookPath, f.error)), -}); +## How to use -await build.build(); +```sh +pnpm add @statewalker/notebook-build @statewalker/webrun-builder @statewalker/webrun-files @statewalker/webrun-modules ``` -`build()` runs to convergence and returns. Call it again to pick up changes. +`@statewalker/webrun-builder`, `@statewalker/webrun-files` and `@statewalker/webrun-modules` are +peer dependencies. There is one entry point, `@statewalker/notebook-build` (ESM). + +Create a build with `newNotebookBuild(options)` and call `build()`. `build()` runs to +convergence and returns; call it again to pick up changes. | option | meaning | | --- | --- | | `notebooks` | Source tree. Scanned recursively; `.md` and `.html` are notebooks, everything else is data. | | `output` | Where pages, attachments and (in static mode) the dependency closure are written. | | `cache` | Where the serialized-notebook artifacts and the incremental state live. | -| `moduleServer` | Resolves and serves npm packages — `@statewalker/webrun-modules`. | +| `moduleServer` | Resolves and serves npm packages (`ModuleServerLike`: `resolve`, `listResources`, `listPackageFiles`, `fetch`). `newModuleServer` from `@statewalker/webrun-modules` provides it. | | `dom` | A `document` and a `DOMParser`. Nothing in this package reads a DOM global. | | `mode` | `"static"` (default) materializes the whole dependency closure into `output`; `"hosted"` leaves it to a live module server. | | `basePath` | URL prefix the module server serves packages under (default `/_m/`). | | `stylesUrl` | Stylesheet the pages link. Unset, the pages carry no styles at all. | -| `onRebuilt` | Called once per converged build with every output path that changed — page, attachments and closure. | -| `onFailed` | Called once per converged build with the notebooks that failed. One bad notebook never stops the others. | +| `onRebuilt` | Called once per converged build with every output path that changed: page, attachments and closure. | +| `onFailed` | Called once per converged build with `NotebookFailure[]` (`{ notebookPath, error }`). One bad notebook never stops the others. | +| `logger` | A `Logger` from `@statewalker/webrun-builder`. Defaults to a no-op logger. | -The three `FilesApi` instances must be three distinct directories. The build writes a hidden -probe file to prove it, because two handles on one directory would feed the build's own output -back in as sources. +## Examples -## Cell modes +### Build a directory of notebooks under Node -`js`, `ts`, `ojs` and `sql` cells are compiled and run in the page; every other mode renders as -prose. +```ts +import { newNotebookBuild } from "@statewalker/notebook-build"; +import { NodeFilesApi } from "@statewalker/webrun-files-node"; +import { newModuleServer } from "@statewalker/webrun-modules"; +import { JSDOM } from "jsdom"; + +const notebooks = new NodeFilesApi({ rootDir: "./notebooks" }); +const output = new NodeFilesApi({ rootDir: "./.out" }); +const cache = new NodeFilesApi({ rootDir: "./.cache/build" }); +const moduleServer = newModuleServer({ + cache: new NodeFilesApi({ rootDir: "./.cache/modules" }), + basePath: "/_m/", +}); -### SQL cells +const { window } = new JSDOM(""); + +const build = newNotebookBuild({ + notebooks, + output, + cache, + moduleServer, + dom: { document: window.document, parser: new window.DOMParser() }, + mode: "static", + stylesUrl: "/_m/@observablehq/notebook-kit@2.6.6/dist/src/styles/index.css", + onRebuilt: (changed) => console.log("changed", changed), + onFailed: (failures) => failures.forEach((f) => console.error(f.notebookPath, f.error)), +}); -A SQL cell's `database` and `output` attributes are what make it work, and only a notebook-kit -HTML source can carry them — a Markdown fence has no syntax for an attribute. +await build.build(); +``` -`database` picks the mode, and the two are notebook-kit's, not this package's: +In a browser, pass `{ document, parser: new DOMParser() }` and any browser `FilesApi`. -- `database="var:db"` (the default) is **live**: the cell compiles to - ``DatabaseClient.of(db, "db").sql`…` `` and queries whatever the notebook's own `db` variable - is. Any object with a `sql` tagged-template function will do — for instance a - [`@statewalker/notebook-db`](https://www.npmjs.com/package/@statewalker/notebook-db) client - over a `@statewalker/db-api` `Db`. -- `database="warehouse"` is **precomputed**: notebook-kit's client runs no SQL at all. It - fetches `.observable/cache/-.json`, and that path is relative to the PAGE, so - a notebook at `/reports/q3.html` reads `/reports/.observable/cache/…`. Writing those files is - `@statewalker/notebook-db`'s `precomputeQueries`. - -A Markdown ` ```sql ` fence is therefore a cell with NEITHER attribute, and since `ca27d82` -that is a live cell, not prose: it compiles to ``DatabaseClient.of(db, "db").sql`…` `` against -the notebook's own `db` variable (the `database` default), and with no `output` it is -anonymous — it runs and displays its result, and nothing downstream can name its rows. A -notebook that wants to reference the rows, or to query anything other than `db`, needs a -notebook-kit HTML source. Measured against notebook-kit 2.6.4, a bare `sql` cell transpiles to -`inputs: ["DatabaseClient", "db"]`, `outputs: []`, no singular `output` and `autodisplay: -true` — a consumer of a `db` variable the notebook must define elsewhere, and a producer of -nothing. - -`output="revenue"` exposes the cell's rows to the rest of the notebook. It is notebook-kit's -*singular* output — `outputs` stays empty for a SQL cell — and two cells claiming one name fail -the build exactly as two `const x` cells do. - -SQL results render through notebook-kit's default inspector. notebook-kit's own Vite plugin -uses `displayMode: "table"` instead; this package does not, because that display path is -`import("…/stdlib/inputs.js")`, whose first line imports `@observablehq/inputs` from jsDelivr. - -### The modes that stay prose - -Not "unsupported": notebook-kit's `transpile()` returns a real body for each of them. They are -left inert because the body cannot run in a page this build produces. - -| mode | why | +### Use the stages directly + +The stages `newNotebookBuild` is made of are exported too: + +```ts +import { + parseMarkdown, + renderPage, + resolveNotebook, + transpileNotebook, +} from "@statewalker/notebook-build"; + +const nb = parseMarkdown("# Hello\n\n```js\nconst x = 1 + 1\n```\n"); +const pins = await resolveNotebook(nb, { moduleServer }, "/hello.md"); // npm specifier -> URL +const cells = transpileNotebook(nb, pins); +const html = renderPage(nb, cells, { runtimeUrl }); // runtimeUrl: notebook-kit's runtime module URL +``` + +| export | what it does | | --- | --- | -| `html`, `tex`, `dot`, `sql.view` | need the `htl`, `tex`, `dot` and `Inputs` builtins, each of which notebook-kit loads from `cdn.jsdelivr.net`. A static export that reaches a CDN is not a static export. | -| `node`, `python`, `r` | are data-loader cells: `Interpreter(…).run(src)` fetches `.observable/cache/.bin`, an artifact a build-time interpreter stage produces. This build has none, so every such cell would 404 — worse than rendering inert. | -| `md` | is rendered at build time with markdown-it, into the document body, so prose is readable with JavaScript off. | +| `parseMarkdown(source)` | Markdown source to a notebook-kit `Notebook`. | +| `parseNotebookHtml(html, dom)` | notebook-kit HTML document to a `Notebook`. | +| `serializeNotebook(nb, dom)`, `notebookHash(html)` | Serialize a `Notebook` to notebook-kit HTML; hash that HTML. | +| `collectSpecifiers(nb)`, `isNpmSpecifier(s)`, `toModuleRef(s)` | Find a notebook's import specifiers and turn `npm:` ones into module refs. | +| `resolveNotebook(nb, { moduleServer }, notebookPath)` | Resolve every npm import to a pinned URL (`PinMap`). Throws `ResolveError`. | +| `transpileNotebook(nb, pins)` | Compile the code cells to `CellDefinition[]`. | +| `renderPage(nb, cells, { runtimeUrl, stylesUrl })` | Render the page HTML. | +| `copyAttachments(nb, source, output, notebookPath)` | Copy the `FileAttachment`s a notebook references; returns `CopiedAttachment[]` (path and hash). | +| `materializeDeps(pins, server, output, basePath)`, `ASSET_EXTENSIONS` | Write the static dependency closure into `output`. | + +## Internals -## What is published +### What a notebook publishes For `/reports/q3.md`: -- `/reports/q3.html` — the page: one root element per cell, one `define()` per code cell. -- every `FileAttachment("…")` it references, copied to the same relative path — including one - inside a `sql` cell's `${…}` interpolation, which notebook-kit compiles as JavaScript like - any other cell's. An attachment that resolves outside the notebook's own directory is - refused. +- `/reports/q3.html`: the page, with one root element per cell and one `define()` per code cell. +- every `FileAttachment("…")` it references, copied to the same relative path, including one + inside a `sql` cell's `${…}` interpolation, which notebook-kit compiles as JavaScript like any + other cell's. - in static mode, the dependency closure under `basePath`: every JS-reachable module, plus two - kinds of file the JS graph never imports and a closure built from it alone would therefore - miss (such an export looks perfect and dies at the first wasm instantiation): - - the `.wasm`, `.css` and font files a package ships; - - the classic worker scripts it ships (`*.worker.js`, and not their `.map` siblings). These - are fetched with `?raw` so the module server's CJS→ESM transform cannot wrap them — - duckdb-wasm's worker bundles are UMD, and a classic worker cannot parse the `import`/ - `export` a wrapped one would contain. The `?raw` is on the fetch only: the file is written - at its plain `.worker.js` path, so the site serves it with a JavaScript content type, which - is what the spec requires of a classic worker script. + kinds of file the JS graph never imports. Without them the export looks complete and fails at + the first wasm instantiation: + - the `.wasm`, `.css` and font files a package ships (`ASSET_EXTENSIONS`); + - the classic worker scripts it ships (`*.worker.js`, not their `.map` files). These are + fetched with `?raw` so the module server's CJS-to-ESM transform does not wrap them: DuckDB's + worker bundles are UMD, and a classic worker cannot parse the `import`/`export` a wrapped + one would contain. The file is written at its plain `.worker.js` path, so the site serves it + with a JavaScript content type, as the spec requires for a classic worker script. Deleting a notebook prunes exactly what it published, minus anything another notebook still claims. -## Incrementality +### Why the three stores must be three directories + +The engine scans `notebooks`, the serialized artifacts go to `cache`, and pages go to `output`. +If two of them were the same directory, the scanner would find the files the build just wrote +and feed them back in as sources, and the site would publish the build's own cache. The +constructor rejects one instance passed twice; the first `build()` also writes a hidden probe +file into `cache` and `output` and looks for it in the others, because two `FilesApi` instances +can point at one directory. Overlap deeper down (an `output` rooted inside the notebooks tree) +is not detected. + +### What makes a page rebuild + +A notebook is re-derived unless everything its published page depended on is unchanged: the +serialized notebook (compared by content hash, so a touched file with identical bytes is +reused), the build configuration (`mode`, `basePath`, `stylesUrl`), the resolved pin map, the +content of every attachment, and the presence of every output. A notebook with no recorded +successful build is retried on the next run without its source being touched. + +### Cell modes + +`js`, `ts`, `ojs` and `sql` cells are compiled and run in the page. The other modes render as +inert prose, because their compiled body cannot run in a page this build produces: + +| mode | why it stays prose | +| --- | --- | +| `html`, `tex`, `dot`, `sql.view` | They need the `htl`, `tex`, `dot` and `Inputs` builtins, which notebook-kit loads from `cdn.jsdelivr.net`. A static export must not depend on a CDN. | +| `node`, `python`, `r` | Data-loader cells: `Interpreter(…).run(src)` fetches `.observable/cache/.bin`, produced by a build-time interpreter stage this build does not have. Every such cell would 404. | +| `md` | Rendered at build time with markdown-it into the document body, so prose is readable with JavaScript off. | + +### SQL cells + +A SQL cell's `database` and `output` attributes decide how it runs, and only a notebook-kit HTML +source can carry them; a Markdown fence has no attribute syntax. + +- `database="var:db"` (the default) is **live**: the cell compiles to + ``DatabaseClient.of(db, "db").sql`…` `` and queries the notebook's own `db` variable. Any + object with a `sql` tagged-template function works, for example a `@statewalker/notebook-db` + client over a `@statewalker/db-api` `Db`. +- `database="warehouse"` is **precomputed**: notebook-kit's client runs no SQL. It fetches + `.observable/cache/-.json` relative to the page, so a notebook at + `/reports/q3.html` reads `/reports/.observable/cache/…`. `precomputeQueries` from + `@statewalker/notebook-db` writes those files. +- `output="revenue"` exposes the cell's rows to the rest of the notebook. Two cells declaring one + name fail the build, as two `const x` cells do. + +A Markdown ` ```sql ` fence has neither attribute, so it is a live, anonymous cell: it queries +`db`, displays its result, and nothing downstream can name its rows. The notebook must define +`db` in another cell. + +SQL results render through notebook-kit's default inspector, not `displayMode: "table"`: the +table display imports `@observablehq/inputs` from jsDelivr. + +### What this package does not do with databases + +`NotebookBuildOptions` has no database option, and nothing here turns a parsed notebook into +`PrecomputeRequest`s. A `database="warehouse"` cell's cache file is therefore never written by +this build; the caller must assemble the requests and run `precomputeQueries` itself. Only live +SQL cells work from a build alone. + +### Failures you will see + +Each failure is reported through `onFailed` for its notebook; the other notebooks still build. -A notebook is re-derived unless everything the published page depended on is unchanged: the -serialized notebook, the build configuration (`mode`, `basePath`, `stylesUrl`), the resolved -pin map, the content of every attachment, and the presence of every output. A notebook with no -recorded successful build — one that failed, transiently or not — is retried on the next run -without needing its source touched. +- `` newNotebookBuild: `notebooks`, `output` and `cache` must be distinct FilesApi instances `` (thrown by the constructor) +- `` newNotebookBuild: `cache` and `notebooks` are the same directory — they must be distinct FilesApi instances over distinct roots `` +- `/a.md: /a.md and /a.html all publish to /a.html — rename all but one` +- `/a.md: cannot resolve import "npm:…": …` (`ResolveError`) +- `/a.md: attachment "…" resolves outside the notebook's own directory (/)` +- `/a.md: cannot find attachment "data.csv" (expected at /data.csv)` +- `/a.md: "x" is declared by two cells (cell 1 and cell 3)` -Two sources that would publish to the same page (`report.md` and `report.html`) are both -reported as failures rather than one silently overwriting the other. +### Dependencies -## Status +- `@observablehq/notebook-kit`: parsing, serialization, transpilation and the page runtime. +- `markdown-it`: renders prose cells at build time. +- `@statewalker/webrun-builder` (peer): the incremental build engine. +- `@statewalker/webrun-files` (peer): the storage interface. +- `@statewalker/webrun-modules` (peer): the module server the build expects. -`html`, `tex`, `dot`, `sql.view`, `node`, `python` and `r` cells parse and render as inert -prose — see "Cell modes" above for why each one is left out rather than wired up. +## License -This package has NO build-time database wiring of any kind — not a stub, not a placeholder. -`NotebookBuildOptions` has no `databases` option, and nothing here derives a -`PrecomputeRequest` from a parsed notebook, so a `database="warehouse"` cell's -`.observable/cache/…json` is never written by this build. `@statewalker/notebook-db` exports -`precomputeQueries`, which writes exactly those files, but a caller must assemble the requests -and run it itself. Only the live path (`database="var:db"`, the Markdown fence default) works -end to end from a build alone. +MIT diff --git a/packages/notebook-db/README.md b/packages/notebook-db/README.md index ab117b1..cf1d4e7 100644 --- a/packages/notebook-db/README.md +++ b/packages/notebook-db/README.md @@ -1,21 +1,48 @@ # @statewalker/notebook-db -Back a notebook-kit SQL cell with any [`@statewalker/db-api`](https://github.com/statewalker/statewalker-db) -`Db` — DuckDB, SQLite, or whatever else implements the interface. +## What it is -notebook-kit's `DatabaseClient.of(source, name)` accepts any object exposing a `sql` tagged-template -function as a database source; it never has to know about db-api. This package supplies that -object. +Backs notebook-kit SQL cells with any `@statewalker/db-api` `Db`: DuckDB, SQLite, or anything +else that implements the interface. It provides a live client for a page, a registry that opens +databases by name, and a build-time step that writes the JSON files notebook-kit's precomputed +SQL cells fetch. + +## Why it exists + +notebook-kit's `DatabaseClient.of(source, name)` accepts any object with a `sql` +tagged-template function as a database. This package supplies that object over a db-api `Db`, +so notebook-kit needs to know nothing about db-api and a notebook can use any db-api driver. +For precomputed cells notebook-kit runs no SQL at all; it fetches a cache file at a path derived +from hashes it does not export. `precomputeQueries` writes those files at the paths and in the +format notebook-kit reads. + +## How to use ```sh -npm install @statewalker/notebook-db +pnpm add @statewalker/notebook-db @statewalker/db-api @statewalker/webrun-files ``` +`@statewalker/db-api` and `@statewalker/webrun-files` are peer dependencies. Bring a db-api +driver as well, for example `@statewalker/db-duckdb-node` or `@statewalker/db-duckdb-browser`. +There is one entry point, `@statewalker/notebook-db` (ESM). It runs in the browser and under +Node; which one depends on the driver. + +| export | what it does | +| --- | --- | +| `newDbClient(db)` | Wraps a db-api `Db` as a `NotebookDbClient` (`sql`, `query`, `close`). | +| `newLiveDatabases({ open })` | A registry that opens databases by name on first use (`get`, `closeAll`). | +| `precomputeQueries(requests, databases, output)` | Runs `PrecomputeRequest`s at build time and writes notebook-kit's cache files into a `FilesApi`. Returns the paths written, one per request. | +| `cachePathFor(notebook, database, strings, params)` | The site path of the cache file for one query on one page. | + +## Examples + +### A client over a db-api database + ```ts import { newDbClient } from "@statewalker/notebook-db"; -import { newDuckDb } from "@statewalker/db-duckdb-node"; // or any other db-api driver +import { newNodeDuckDb } from "@statewalker/db-duckdb-node"; // or any other db-api driver -const db = await newDuckDb(); +const db = await newNodeDuckDb(); const client = newDbClient(db); // Tagged-template form, the shape a notebook SQL cell compiles to: @@ -27,49 +54,77 @@ const rows2 = await client.query("SELECT * FROM t WHERE id = ?", [id]); await client.close(); ``` -## Live databases in a page +### Live databases in a page -`newLiveDatabases` is the registry a built notebook page wires a **live** SQL cell to. A SQL -cell whose `database` is `var:db` compiles to ``DatabaseClient.of(db, "db").sql`…` ``, so the -notebook needs a `db` variable holding something with a `sql` tagged template — which is what -`get(name)` resolves to. +A live SQL cell (`database="var:db"`) compiles to ``DatabaseClient.of(db, "db").sql`…` ``, so +the notebook needs a `db` variable holding something with a `sql` tagged template. `get(name)` +resolves to such a client. ```ts import { newLiveDatabases } from "@statewalker/notebook-db"; import { newBrowserDuckDb } from "@statewalker/db-duckdb-browser"; -const databases = newLiveDatabases({ open: () => newBrowserDuckDb({ bundles }) }); +const databases = newLiveDatabases({ open: (name) => newBrowserDuckDb({ bundles }) }); // in the notebook's own js cell: const db = await databases.get("warehouse"); ``` -A database is opened on first use and reused after; two callers racing the first `get` share one -`open`. A *rejected* open is not cached — OPFS can be unavailable and a wasm bundle can be -blocked, and both are transient — and the rejection is rewrapped with the database name, because -"OPFS unavailable" on its own does not say which cell to fix. `closeAll()` closes every database -opened so far and empties the registry, so a second teardown does not close them twice and a name -requested afterwards gets a fresh database rather than a closed one. +### Precompute a query at build time + +```ts +import { newDbClient, precomputeQueries } from "@statewalker/notebook-db"; + +const written = await precomputeQueries( + [ + { + notebook: "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/reports/q3.html", // the page the cell lives on + database: "warehouse", + strings: ["SELECT * FROM sales WHERE year = ", ""], + params: [2026], + }, + ], + new Map([["warehouse", newDbClient(db)]]), + output, // FilesApi of the built site +); +// written[0] is "/reports/.observable/cache/-.json" +``` + +`@statewalker/notebook-build` does not create these requests; the caller assembles them. + +## Internals + +### Interpolations are bound, never concatenated + +In the tagged template, the SQL text is `strings.join("?")` and the interpolated values are +passed as a separate `params` array to `Db.query(sql, params)`. A cell that interpolates +user-supplied data cannot inject SQL. + +### Why a live database is opened once, and a failed open is retried + +A database is opened on first `get` and reused; two callers racing the first `get` share one +`open`. A rejected open is not cached: OPFS can be unavailable and a wasm bundle can be blocked, +and both can be transient. The rejection is rewrapped with the name, because the driver's message +alone does not say which cell to fix: -## Parameter binding +``` +cannot open database "warehouse": +``` -Interpolations in the tagged template become bound parameters — `strings.join("?")` for the SQL -text, the interpolated values as a separate `params` array — passed to `Db.query(sql, params)`. -They are never concatenated into the SQL string. A cell querying user-supplied data would -otherwise be an injection. +`closeAll()` closes every database opened so far and empties the registry, so a second teardown +does not close them twice and a later `get` opens a fresh database rather than returning a closed +one. -## BIGINT columns +### BIGINT becomes number, or fails loudly -DuckDB answers `count(*)`, an integer `sum()` and any `BIGINT` column with a JavaScript -`bigint`. `newDbClient` converts those to `number`, in both `sql` and `query`, and so does the -precompute stage before it serializes. That is the same conversion notebook-kit's own -`DatabaseClient.revive` performs (`row[name] = Number(value)`), so a live cell and a -precomputed one show the same value; a string would make `count(*)` render as `"3"` and -`rows[0].c + 1` produce `"31"`. +DuckDB returns `count(*)`, an integer `sum()` and any `BIGINT` column as a JavaScript `bigint`. +`newDbClient` converts those to `number` in both `sql` and `query`, and the precompute step does +the same before it serializes. This matches notebook-kit's own `DatabaseClient.revive` +(`row[name] = Number(value)`), so a live cell and a precomputed one show the same value. A string +would make `count(*)` render as `"3"` and `rows[0].c + 1` produce `"31"`. -A double holds every integer below 2^53 exactly and none above it, so a value outside -`Number.MAX_SAFE_INTEGER` is an ERROR naming the column and the exact value, never a silently -rounded result: +A double holds every integer up to 2^53 exactly and none above it, so a value outside +`Number.MAX_SAFE_INTEGER` is an error naming the column and the value, never a rounded result: ``` cannot represent BIGINT column "id" as a JSON number: 9007199254740993 is outside @@ -77,41 +132,55 @@ Number.MAX_SAFE_INTEGER (9007199254740991) and would become 9007199254740992. Cast it in SQL (for example `CAST("id" AS VARCHAR)`) to keep the exact value. ``` -Nested `LIST` and `STRUCT` values are converted too; `Date`, `Uint8Array` and anything else -carrying its own prototype is left untouched. - -## The precomputed cache file - -`precomputeQueries` writes each query's result at the page-relative path notebook-kit's -`DatabaseClient.sql()` fetches, and writes it as the `{rows, schema}` envelope that client -expects — never a bare array. `sql()` is `fetch(path).then(r => r.json()).then(revive)`, and -`revive` destructures `{rows, schema, ...}` and iterates `schema`, so an array throws -`TypeError: schema is not iterable` on every precomputed cell. - -`schema` here is a **revival directive**, not SQL column-type metadata: db-api reports no SQL -types, and this field only says which values need reconstructing after a JSON round trip. It is -derived by inspecting the values being serialized, and it matters for exactly one case — -`revive` branches on `"bigint"` and `"date"` and ignores every other type. A `Date` column must -be marked `"date"` or the precomputed page gets the ISO string where the live page gets a -`Date`. - -Every row is scanned, not just the first, so a `Date` column whose first row is `NULL` is still -found. A column that is `NULL` in every row is marked `"other"` (nothing observed, nothing to -revive). A column is marked `"date"` only if *every* non-null value is a `Date`: `revive` marks -a column rather than a value, so marking a mixed column would turn its non-dates into -`Invalid Date`. Mixed columns therefore keep their JSON values verbatim, which means a `Date` -inside one arrives as a string — a limit of the cache format, and not reachable from a typed -SQL column. - -`"bigint"` is never emitted: the conversion above already ran, so those columns hold numbers by -the time they are serialized. - -## What this package does not do - -It does not implement SQL composition, dialect-specific identifier quoting, or view/CTE flattening. -Those are notebook-kit's own `sql` tagged template, `SqlFragment`, `SqlView` and `sql.ident` — import -them from `@observablehq/notebook-kit` if a notebook needs to compose fragments before handing them -to a `client`. +Nested `LIST` and `STRUCT` values are converted too; `Date`, `Uint8Array` and other values with +their own prototype are left as they are. Only object rows are converted. A custom client that +returns tuple rows (`[[1n, 2n]]`) passes through, and `JSON.stringify` in `precomputeQueries` +then fails with `Do not know how to serialize a BigInt`. + +### The cache file is `{rows, schema}`, and where it goes depends on the page + +notebook-kit's `DatabaseClient.sql()` is `fetch(path).then(r => r.json()).then(revive)`, and +`revive` destructures `{rows, schema}` and iterates `schema`. A bare array would fail in the page +with `TypeError: schema is not iterable`, so `precomputeQueries` always writes the envelope. + +The fetched path, `.observable/cache/-.json`, has no leading slash, so the +browser resolves it against the page's directory. That is why every `PrecomputeRequest` names +its page (`notebook`). A query run by pages in two directories is executed once and written +twice, once per directory. The hash covers `JSON.stringify([strings, ...params])`, so two queries +that differ only in a bound value, or in whether a literal is inline or bound, get different +files. `hash` and `nameHash` are not exported by notebook-kit, so this package reimplements them, +and its tests compare the result with notebook-kit's own `cachePath`. + +A request whose `database` has no client in the map fails: + +``` +precomputeQueries: no database configured for "warehouse" (query: SELECT * FROM sales WHERE year = ?) +``` + +### `schema` says what to revive, not the SQL types + +db-api reports no SQL types. `schema` only tells `revive` which values to rebuild after the JSON +round trip, and `revive` acts only on `"bigint"` and `"date"`. It is derived from the values: + +- Every row is scanned, so a `Date` column whose first row is `NULL` is still found. +- A column that is `NULL` in every row is `"other"`. +- A column is `"date"` only if every non-null value is a `Date`. `revive` converts whole columns, + so marking a mixed column would turn its non-dates into `Invalid Date`. A `Date` inside a mixed + column therefore arrives as a string. +- `"bigint"` is never written: those values are already numbers. + +### What this package does not do + +It does not compose SQL, quote identifiers for a dialect, or flatten views and CTEs. For that use +notebook-kit's own `sql` tagged template, `SqlFragment`, `SqlView` and `sql.ident` from +`@observablehq/notebook-kit`, and hand the result to a client. + +### Dependencies + +- `@statewalker/db-api` (peer): the `Db` interface. No driver is required by the package itself. +- `@statewalker/webrun-files` (peer): path helpers and `FilesApi` for writing cache files. +- No runtime dependency on `@observablehq/notebook-kit`: the only use is a type, which is inlined + in the built `.d.ts`. ## License diff --git a/packages/notebook-site/README.md b/packages/notebook-site/README.md index 7e253e6..054086b 100644 --- a/packages/notebook-site/README.md +++ b/packages/notebook-site/README.md @@ -1,98 +1,132 @@ # @statewalker/notebook-site -Composes one [`SiteHandler`](https://github.com/statewalker/webrun-wire) that serves a site built -by [`@statewalker/notebook-build`](../notebook-build): pages and attachments from a `FilesApi`, -module dependencies from a live module server, and the rebuild stream from -[`@statewalker/webrun-http-events`](../notebook-events). +## What it is -It owns the composition and nothing else — no routing of its own, no build, no transport. The -same handler runs under Node, in a Worker, and behind a browser ServiceWorker. +Composes one `SiteHandler` (a `Request` to `Response` function from +`@statewalker/webrun-site-builder`) that serves a site built by `@statewalker/notebook-build`: +pages and attachments from a `FilesApi`, npm module dependencies from a live module server, and +rebuild notifications as a Server-Sent Events stream from `@statewalker/webrun-http-events`. -```sh -npm install @statewalker/notebook-site -``` +## Why it exists -## Usage +A built notebook site needs three things mounted side by side, with rules that are easy to get +wrong: module requests must not shadow pages, a static export must serve its modules from files, +paths must be percent-decoded, and nothing may throw out of the handler. This package owns that +composition and nothing else: no routing of its own, no build, no transport. The same handler +runs under Node, in a Worker, and behind a browser ServiceWorker. -```ts -import { newNotebookSite, primeModules } from "@statewalker/notebook-site"; - -const handler = newNotebookSite({ - output, // FilesApi: what notebook-build wrote - moduleServer, // hosted mode only — omit for a static export - events, // PubSub from @statewalker/webrun-http-events — omit for no event stream - basePath: "/_m/", - eventsPath: "/_events", -}); +## How to use -const response = await handler(new Request("http://localhost/reports/q3.html")); +```sh +pnpm add @statewalker/notebook-site @statewalker/webrun-files @statewalker/webrun-site-builder ``` +`@statewalker/webrun-files` and `@statewalker/webrun-site-builder` are peer dependencies. There +is one entry point, `@statewalker/notebook-site` (ESM). It exports `newNotebookSite` and +`primeModules`. + | option | meaning | | --- | --- | | `output` | The built site: pages, attachments, and in static mode the materialized dependency closure. | -| `moduleServer` | Serves `basePath/*` in hosted mode. **Omit it for a static export** — then nothing claims that prefix and the request falls through to `output`. | +| `moduleServer` | Serves `basePath/*` in hosted mode. Omit it for a static export: then nothing claims that prefix and the request falls through to `output`. | | `basePath` | Where module dependencies are mounted (default `/_m/`). Must match the build's `basePath`. | -| `events` | Rebuild notifications. Omit to serve without an event stream. | +| `events` | A `PubSub` for rebuild notifications. Omit to serve without an event stream. | | `eventsPath` | Where the event stream is mounted (default `/_events`). | -| `directoryIndex` | File served for a directory request (default `index.html`). `webrun-site-builder` has no default here, and without one a directory request 404s, which reads as a broken build. | +| `directoryIndex` | File served for a directory request (default `index.html`). `webrun-site-builder` has no default, and without one a directory request returns 404. | -## Dispatch +## Examples -Endpoints are matched before files, so both mounts must stay clear of notebook paths — hence the -underscore prefixes. +### Serve a built site -- `basePath` and `eventsPath` are prefixes of path **segments**, not of the string. `/_m/*` does - not match `/_module-notes.html`, and `/_events/*` does not match `/_eventsource-guide.html`. -- A trailing slash is optional and means nothing: `"/_m/"` and `"/_m"` mount the same place. The - two defaults are spelled inconsistently — `/_m/` has one, `/_events` does not — so either - spelling of either option has to work, and does. -- The site root is refused. `basePath: "/"` would make the module server claim every page, - `/index.html` included, so it throws rather than serving nothing. +```ts +import { newNotebookSite } from "@statewalker/notebook-site"; +import { newPubSub } from "@statewalker/webrun-http-events"; -## Paths are percent-decoded +const events = newPubSub(); -A `SiteHandler` receives a `Request`, and a URL pathname is percent-encoded by definition. Nothing -below this package decodes it, so `newNotebookSite` does — otherwise `/My%20Notebook.html` would -404 here while the identical static export, served by any ordinary HTTP server, worked. +const handler = newNotebookSite({ + output, // FilesApi: what notebook-build wrote + moduleServer, // hosted mode only; omit for a static export + events, // omit for no event stream + basePath: "/_m/", + eventsPath: "/_events", +}); -Decoding is per segment. `%2f` is not a path separator to the URL parser, and it is not turned -into one: a segment that decodes to a dot-segment or to anything containing a separator rejects -the path rather than reaching the backend. A `%` that is not a valid escape is left alone, since a -filename may legitimately contain one. +const response = await handler(new Request("http://localhost/reports/q3.html")); +``` -## Priming +### Warm the module cache before serving ```ts +import { primeModules } from "@statewalker/notebook-site"; + +// moduleServer: anything with prime(ref), e.g. newModuleServer from @statewalker/webrun-modules const { primed, failed } = await primeModules(moduleServer, [ { pkg: "d3", version: "7" }, { pkg: "katex", version: "0.16", subpath: "dist/katex.mjs" }, ]); ``` -Lazy emission of `~deps` proxy files races concurrent browser fetches — a cold first load failed -roughly one time in four with a link error naming a proxy that had not finished being written. -`primeModules` warms them first. It is serial (concurrency re-creates the contention it exists to -avoid), it deduplicates refs, and one unresolvable package lands in `failed` instead of leaving -the rest of the site cold. +## Internals + +### How a request is dispatched + +``` +Request ──► basePath/* ──► moduleServer.fetch (only when moduleServer is given) + ──► eventsPath/* ──► events.handler (only when events is given) + ──► /* ──► output, percent-decoded, directoryIndex for directories + any throw ──► logged as "[notebook-site]", answered 500 "internal error" +``` + +Endpoints are matched before files, so both mounts must stay clear of notebook paths; that is +why the defaults start with an underscore. + +- The mounts are prefixes of path segments, not of the string: `/_m/*` does not match + `/_module-notes.html`, and `/_events/*` does not match `/_eventsource-guide.html`. +- A trailing slash is optional: `"/_m/"` and `"/_m"` mount the same place. The two defaults are + spelled differently (`/_m/` with a slash, `/_events` without), so either spelling of either + option works. +- The site root is refused. At `/` the endpoint would claim every page, `/index.html` included, + so `newNotebookSite` throws: + `newNotebookSite: basePath must not be the site root (got "/"); an endpoint mounted there claims every page, including /index.html` + +### Paths are percent-decoded + +A URL pathname is percent-encoded, and nothing below this package decodes it. Without decoding, +`/My%20Notebook.html` would 404 here while the same static export served by an ordinary HTTP +server works. Decoding is per segment. `%2f` is not turned into a path separator: a segment that +decodes to a dot-segment or contains a separator makes the request a 404 instead of reaching the +backend. A `%` that is not a valid escape is left as is, because a file name may contain one. + +### Why nothing escapes the handler + +A throw in any layer is logged and answered with a `500`, because a rejected handler inside a +ServiceWorker breaks every open page. + +This does not cover a failure inside the response body. File responses stream from +`output.read()` lazily, so a backend that fails mid-read fails after the handler has returned: +the client gets a `200` with a truncated body, with no `500` and no log. + +### Why priming is serial -## Errors +The module server writes its `~deps` proxy files lazily, and that races concurrent browser +fetches: a cold first load can fail with a link error naming a proxy that is not written yet. +`primeModules` resolves the refs first. It runs one ref at a time, because concurrent priming +re-creates the contention it exists to avoid. It deduplicates refs, and a package that fails to +resolve lands in `failed` (with its error message) instead of stopping the rest. -Nothing escapes the handler as a rejection: a throw in any layer is logged and answered with a -`500`. A thrown handler in a ServiceWorker takes down every open page. +### Hosting behind a ServiceWorker -**Known gap:** this does not cover a failure *inside the response body*. `newServeFiles` returns a -`Response` whose stream pulls from `output.read()` lazily, so a backend that throws mid-read fails -long after the handler returned — the caller gets a `200` with a truncated body, and there is no -`500` and no log. Closing it needs either a buffered body or a trailer-capable stream. +`HostedSiteBuilder` from `@statewalker/webrun-site-host` runs the handler behind a ServiceWorker, +not inside one: the worker intercepts `fetch` and relays over a `MessagePort`, and the page-side +`SwHttpAdapter` (from `@statewalker/webrun-http-browser/sw`) calls the `SiteHandler`. The package is compiled without the DOM lib. That is a +compile-time guard, not proof that the handler never touches the DOM at run time. -## Hosting behind a ServiceWorker +### Dependencies -`@statewalker/webrun-site-host`'s `HostedSiteBuilder` runs the handler behind a ServiceWorker, not -inside one: the worker half intercepts `fetch` and relays over a `MessagePort`, and the page-side -`SwHttpAdapter` is what invokes the `SiteHandler`. The shipped package is compiled without the DOM -lib, which is a compile-time guard — it is not runtime proof that the handler never touches the -DOM, and nothing here executes in worker scope. +- `@statewalker/webrun-site-builder` (peer): `SiteBuilder`, which does the routing and file serving. +- `@statewalker/webrun-files` (peer): the `FilesApi` the site is served from. +- `@statewalker/webrun-http-events`: the `PubSub` type for the event stream. ## License