Skip to content

Add opt-in persistent conversation checkpoints - #21

Open
owenqwenstarsky wants to merge 1 commit into
Edge0-AI:mainfrom
owenqwenstarsky:feature/persistent-conversation-cache
Open

owenqwenstarsky wants to merge 1 commit into
Edge0-AI:mainfrom
owenqwenstarsky:feature/persistent-conversation-cache

Conversation

@owenqwenstarsky

Copy link
Copy Markdown

Repeated coding conversations currently recompute their full prompt after each request or process restart. This adds opt-in persistent inference checkpoints so the server and Python engine can restore the longest saved token prefix and process only the unmatched suffix.

The implementation uses a radix index backed by SQLite, immutable shared safetensors KV blocks, and checkpoint-specific recurrent, logits, and prerouter state. It includes namespace fingerprints, checksums, atomic publication, cross-process locking, bounded LRU eviction, exact-message tokenization caching, CLI inspection/clearing, and cached-token usage metrics. Caching defaults to disabled; when enabled, the defaults are 20 GiB and a 2,048-token interval.

Validation:

  • 84 regular tests passed; 1 skipped.
  • Both local 8B real-weight tests passed. Two 35B tests were explicitly skipped because the weights were unavailable.
  • A 2,799-token coding fixture produced identical output tokens with caching disabled, on first write, on repeat, and after process restart.
  • On an M5 MacBook Air with 24 GB RAM, measured TTFT was 6.51 s disabled, 6.77 s on first write, 0.50 s on repeat, and 0.61 s after restart. These are sequential single samples with a warm OS filesystem cache, not controlled cold-SSD measurements.

Publication increased peak MLX allocation by about 24% in this sample. Shared blocks are currently reserialized synchronously, so retained-storage deduplication does not eliminate write work. Real-weight 35B validation, short-prefix break-even measurements, and sustained/cold-filesystem trials remain outstanding. This does not add GGUF support, KV quantization, or active-context offloading.

Usage, raw benchmark results, and continuation limitations are documented in docs/conversation-cache.md, docs/conversation-cache-results.md, and docs/benchmarks/conversation-cache-m5.json.


Disclosure: This implementation, tests, documentation, benchmarks, and PR description were produced with OpenAI Codex assistance. Codex also ran the reported local validation and prepared this draft PR at the repository owner's request.

@owenqwenstarsky

Copy link
Copy Markdown
Author

This PR is not ready to be merged. It remains a draft pending further review and validation, including real-weight 35B testing and additional cache performance measurements.

@owenqwenstarsky
owenqwenstarsky marked this pull request as ready for review September 12, 2026 15:41
@owenqwenstarsky

Copy link
Copy Markdown
Author

If you think this is a good thing to implement I will continue working on it a bit more and testing it out more to get it ready for any real work.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant