Summary
Open and page long threads without parsing the whole event log. Today opening a thread, the first event of a non-resident thread, and every history page of a non-resident thread each parse the entire JSONL on the host's main loop.
Current behaviour
read_events parses the full <id>.jsonl (crates/services/src/store.rs); called when a client opens a thread (crates/runtime/src/app/sessions.rs) and on the first event of an unloaded thread (crates/runtime/src/app/events.rs).
- Non-resident history pages re-parse from scratch every time (
crates/runtime/src/app/history.rs).
ReadItemOutput ("Load full output") re-folds the whole log.
- Logs on the maintainer's machine: median 0.8 MB, p99 50 MB, max 193 MB.
Proposal
Depends on #523 and #522.
- Two files per thread.
<id>.jsonl holds sealed, compacted turns (append-only). <id>.tail.jsonl holds the in-progress turn's raw events.
- Turn end: fold the tail, compact it, append to the main file, fsync, commit the index rows, delete the tail.
- Index rows (redb):
(session, turn) → { byte offset, first record, record count } and (session, item_id) → location of the final output.
- Crash recovery: if the main file is longer than its committed length, truncate; if a tail exists, rerun the turn-end step.
- Reads: the first snapshot reads only the last turns; backward pages seek by offset;
ReadItemOutput reads one record. All parsing goes through Host::unblock, never the host loop.
- Lazy build: the index for an existing log is built on first open.
- Export: main file + tail, byte-compatible with today's JSONL.
- Cold logs: re-encode logs untouched for N days (and archived logs) as multi-frame gzip (
flate2, already a dependency), about 256 KB of turns per frame, with frame offsets in the index so paging by turn keeps working. Deleting a thread unlinks its files, freeing space immediately.
- Later: persist extracted search text per session so content search stops folding every log (
crates/runtime/src/app/session_search.rs currently parses all logs, archived included, and keeps the text in memory).
Summary
Open and page long threads without parsing the whole event log. Today opening a thread, the first event of a non-resident thread, and every history page of a non-resident thread each parse the entire JSONL on the host's main loop.
Current behaviour
read_eventsparses the full<id>.jsonl(crates/services/src/store.rs); called when a client opens a thread (crates/runtime/src/app/sessions.rs) and on the first event of an unloaded thread (crates/runtime/src/app/events.rs).crates/runtime/src/app/history.rs).ReadItemOutput("Load full output") re-folds the whole log.Proposal
Depends on #523 and #522.
<id>.jsonlholds sealed, compacted turns (append-only).<id>.tail.jsonlholds the in-progress turn's raw events.(session, turn) → { byte offset, first record, record count }and(session, item_id) → location of the final output.ReadItemOutputreads one record. All parsing goes throughHost::unblock, never the host loop.flate2, already a dependency), about 256 KB of turns per frame, with frame offsets in the index so paging by turn keeps working. Deleting a thread unlinks its files, freeing space immediately.crates/runtime/src/app/session_search.rscurrently parses all logs, archived included, and keeps the text in memory).