Skip to content

feat: add owner-coordinated database compaction - #238

Closed
vinceblock99 wants to merge 1 commit into
feat/atomic-redbfrom
feat/database-compact
Closed

vinceblock99 wants to merge 1 commit into
feat/atomic-redbfrom
feat/database-compact

Conversation

@vinceblock99

Copy link
Copy Markdown
Contributor

Repeated writes can leave reclaimable space in the combined atomic.redb. Add atomic compact [--repository PATH] [--json] to let users explicitly reclaim that space through the existing database owner. The command reports the absolute database path, before/after file sizes, and reclaimed bytes.

This follows the database-size investigation: on a copy of a 66-file repository, explicit compaction reduced the file from 19,664,896 to 14,307,328 bytes, reclaiming 5,357,568 bytes (27.2%) while preserving its data. redb already reuses freed pages and can shrink on ordinary commits; explicit compaction can reclaim additional space by reorganizing pages. This PR addresses that measured opportunity without changing the graph/index representation or deleting change objects.

Compaction requires exclusive access, so it is an explicit owner operation rather than maintenance on every command or shutdown:

  • The owner blocks new store leases, drains active requests, and acquires redb's exclusive file lock. Existing reader/writer handles cause a bounded wait using ATOMIC_DB_LOCK_WAIT_MS, then a clear error if still busy. Once started, compaction runs to completion.
  • The database opener lives in the owner's private maintenance module (pub(super)), avoiding a new public Repository opening API. The CLI sends an IPC request; the owner remains available afterward.
  • Sandbox invocations resolve to their canonical database. An additive capability check tells users to restart an older owner before retrying.
  • Size reporting includes redb's allocator-metadata write on close, so it reflects the final file size. A repeat compaction may succeed with zero bytes reclaimed.

Graph, views, provenance, .change files, and unrecorded working files are preserved. This command only operates on an existing atomic.redb: it does not create tables/databases, migrate legacy layouts, delete persistent savepoints, or run automatically. Large repositories should be compacted during a quiet period because other database operations may exhaust their lock-wait budgets while compaction runs.

This Draft PR is stacked on #230 and targets feat/atomic-redb, so its diff contains only compaction and its tests/documentation. Retarget to dev after #230 lands. This PR does not modify atomic-storage.

Validation on macOS:

  • cargo test --workspace --locked --offline: 9,125 passed, 0 failed, 243 ignored, including all 28 owner integration tests.
  • Seven new unit tests and two new process integration tests cover actual reclamation, retained data, missing databases, busy reader/writer handles, future schema rejection, persistent savepoint preservation, owner capability compatibility, lease draining, sandbox resolution, and timeout/retry.
  • The 66-file CLI test preserved all working-file and change-file SHA256 hashes, history, views, and diff output. Restoring a modified recorded file succeeded; a second compact reclaimed zero bytes. The original repository was unchanged.
  • cargo clippy --workspace --locked --offline -- -D warnings, cargo fmt --all -- --check, and the CI semantic diff harness (23/23 assertions) passed. Linux and Windows verification is left to CI.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant