Skip to content

Remote sandboxes over the owner protocol, and signing with a key held elsewhere - #224

Closed
geekgonecrazy wants to merge 25 commits into
devfrom
feat/remote-sandboxes
Closed

geekgonecrazy wants to merge 25 commits into
devfrom
feat/remote-sandboxes

Conversation

@geekgonecrazy

@geekgonecrazy geekgonecrazy commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Lets an agent work in a sandbox that mounts nothing from the host: a tree materialized from the repository's owner, recording through that owner. It also lets a signer keep its key outside the sandbox.

Remote sandboxes

A remote sandbox is a tree of a view, materialized from the repository's owner over iroh. It speaks the existing database-owner protocol; there is no second protocol. Changes recorded in the sandbox land in the repository through the owner.

  • Reach a repository's owner from a remote sandbox, over the same protocol. The owner's iroh endpoints can bind IPv4 only.
  • Record in the sandbox's cache exactly as on the repository, and land the change through the owner. Remote sandboxes live inside Repository, so any caller records, and provenance lands.
  • Read a sandbox's recorded state: diff, restore, log. A sandbox's pointer file is never treated as untracked.
  • A sandbox's cache takes its view's vault. Draft views render as themselves, and their vault is indexed.
  • Render a view's tree without touching disk, and stage and seal a view as it renders.

Signing with a key held elsewhere

  • Attestations can be signed with a key held elsewhere. atomic intent attest --prepare prints the document and the bytes to sign. --signed <file> records an attestation made from them, checked against --identity's key and the intent as it is now.
  • Tokens name the server they are for, as aud.

Merged dev

dev is merged in, including Ed25519 change signing (#214). In atomic-agent/src/identity.rs the merge takes dev's version as-is. That supersedes "fix(agent): never record agent work under the human's key" from this branch: dev falls back to the default signer on purpose (a_human_identity_selected_falls_back_to_the_default_signer).

Testing

  • The remote-sandbox integration tests (atomic-cli/tests/remote_sandbox_integration_test.rs, atomic-repository/tests/remote_cache_parity_test.rs) cover recording and landing through the owner.
  • cargo test -p atomic-agent --lib passes after the merge.
  • cargo test --workspace passes on the merged branch: 57 test binaries, 9075 tests, 0 failures.

`attest_value` splits into `prepare_attestation` (author, content hash and
the exact bytes to sign) and `attach_proof`, so a signer that keeps its key
outside this process — a browser's WebCrypto, a hardware token — produces
the same `eddsa-jcs-2022` attestation the CLI does. `attest_value` is now
those two halves around a local `Signer`.

`atomic-canonical-wasm` exposes that to the browser: prepare, attach and
verify attestations, plus the DID and canonical JSON of a document, all
over JSON strings and byte arrays. `build.sh` builds it with wasm-bindgen;
`smoke.mjs` signs with a non-extractable WebCrypto key, verifies, and
checks tampering is caught.
`atomic intent attest --prepare` prints the document an attestation
signs and the exact bytes to sign, needing only the identity's public
key. `--signed <file>` records an attestation a key holder produced from
it: it must be signed by `--identity`'s key and attest the intent as it
is now, so a signature over a stale or altered intent is refused.

This lets a sandbox that holds only an agent's public identity attest as
that agent, with the key kept by a signing service outside it.
When an agent identity was named (option or ATOMIC_AGENT_IDENTITY) but
couldn't be used — missing, not an agent, unreadable — recording fell
back to the plus-tag author on the human's default key, putting a
person's key on work an agent did, exactly when the caller said it was
an agent's. It now records unkeyed instead, and says so in the log.
Self-signed request tokens now carry `aud`: the bare server URL they
were minted for, normalised. A server that checks it refuses a token
minted for somewhere else, so one leaked from one server can't be
replayed at another within its five minutes. Verifiers that ignore
unknown claims are unaffected.
`Repository::materialize_view_entries` hands each entry of a view — path,
inode, kind, mode, bytes, content hash, conflict-marker line — to a
sink, in path order with directories first, plus the view's Merkle state.
Read-only: no working tree, stat cache or conflict state is written, so
it runs beside other readers. It serves a remote sandbox its tree and
baseline.

The per-file rendering is now one function, `render_view_file`, shared
with `materialize_parallel`, which keeps writing to disk as before.
`.atomic-sandbox` showed up as untracked inside a sandbox, so
`record --all` would record it; a remote pointer carries a token.
Ignore it, and the sandbox's local cache `.atomic-sandbox.d`, like
`.atomic`.
Remote sandboxes use the database owner's protocol rather than a second
one: iroh is another transport beside the local socket, answered by the
same handler. Each remote frame carries a view-scoped sandbox token
(in memory only, hashed); a remote caller gets Ping, the provenance
requests and Materialize, each checked against its token — its own view,
the sessions and turns it started, checkpoints only for changes its view
can see. Opening, renewing and closing sandboxes, and shutdown, stay
local.

The client picks its route from .atomic-sandbox, so provenance from a
remote sandbox goes through the same OwnerJournalSink as a local hook.

`atomic sandbox create --remote` mints the token and writes the
pointer (the owner's iroh address, the view, the token; 0600);
`materialize` writes the view's tree inside the sandbox; `renew` and
`close` manage the token. The owner's iroh key persists in .atomic so
pointers survive a restart.
A change carries the repository's internal node ids and inode numbers,
so a sandbox that records without the repository must read the
repository's own rows. atomic-core's pristine::slice exports and imports
them byte for byte: a skeleton (the view's tree, inodes and positions,
directories, ids of every visible change) and, per record, a graph slice
for the inodes it touches (each file's content and name chain, one hop of
neighbours so find_block resolves as it does on the repository, its CRDT
rows and conflicts). The view is imported flattened, with the
repository's own Merkle state. Inodes the cache adds itself start at
2^62, clear of the repository's.

The change store can hold content spans without their change files, and
consults them first. Repository::{export,import}_sandbox_{skeleton,slice}
wrap both sides; sandbox_slice_inodes is what a cache asks for — changed
files and the directories on the way to changed or new paths.

status: a tracked file missing from disk is "already deleted" only when
the graph holds its vertex; a cache has the tree before the graph.

The parity test records the same edit on a repository and on a cache
built from it — modify, deletions spanning changes, delete, add (existing
and new directory), move, move and edit, and a second record after the
first lands — and requires identical hashes and bytes.
Two more requests on the owner protocol, checked against the sandbox's
token like the rest: FileStates hands the cache the repository's rows and
content for the inodes record will read (only inodes on the view), and
SubmitChange takes the V3 bytes the cache recorded.

The owner applies a submitted change only if it is what it claims and
names nothing outside the view: the bytes hash to the claimed hash; the
view is still at the state it was recorded against; every change it
depends on or refers to that the repository knows is visible on the view;
every node its file operations name is a visible change's
(FileOps::referenced_node_ids); no path touches .atomic, .atomic-sandbox
or .atomic-sandbox.d; and it is new. Submissions go in one at a time.
Refusals write nothing.

In the sandbox, materialize also builds the cache: an ordinary
repository under .atomic-sandbox.d/cache whose working tree is the
sandbox, so a remote pointer opens it and status, log and diff work
locally. record hydrates first, records without saving or applying,
submits, then takes the view's new skeleton and marks the tree clean.
When a change lands, a path the sandbox added gets the repository's
inode.
The cache holds a view's rows, not its content. diff and restore now
load the content they compare against first (FileStates, as record
does), and log fetches the change files the view has and the cache lacks
— a new Changes request, answered only for changes on the token's view,
each file kept only if its bytes hash to what it claims.
…e lands

The remote path moves out of individual commands into Repository, so
every caller gets it — the atomic CLI and an agent's turn hooks alike. A
process installs one RemoteSandboxLink (the atomic binary installs the
owner client at startup); in a remote sandbox's cache, record hydrates
first, write_recorded submits instead of applying (a change that didn't
land fails the record), diff and restore hydrate, log fetches change
files, and publish_provenance_checkpoint publishes the checkpoint in the
repository.

PublishProvenance, the owner's side: the sandbox's own session, and
every change the provenance explains on its view; the graph arrives
serialized, so the repository publishes what the sandbox hashed.

Two authorization fixes the agent flow showed: asking after a session
nobody has written to is allowed (the store answers "no such turn"),
and BindCheckpointHash names a provenance graph, not a change.

An agent's session-start, turn-start, turn-end and session-end in a
remote sandbox now leave its change on the view and its session ledger
and provenance in the repository.
The vault arrives as files, as after a pull, so building the cache
bootstraps the vault tables from them; intents written in the sandbox
are then recorded and land on the view like any other change.
A draft's structural changes (files it adds, moves, deletes) wait in the
deferred tree journal until the view is checked out; TREE is the
checked-out view's. Rendering any other view — materialize_view_entries,
so Materialize and the skeleton after a submit — now projects that
view's journal in a write transaction that is thrown away, and the
skeleton's tree is the view as rendered (implicit directories, inode 0,
left out). Before, a file added on a draft view never appeared
on it.

When a sandbox's change lands, the vault files it touched are indexed
into the repository's vault from the view's content, as a pull does from
the files it writes (vault_record_files, factored out of the working-copy
deflate). An intent created on a draft view is in `intent list` again.

Tests: record parity and a submitted change on a draft, a second record
building on the first, and the agent flow on a draft view off dev.
materialize_view_to (behind `sandbox stage` and `seal`) listed files
from TREE, which is the checked-out view's: a file another view added
was missing from its image. It now writes materialize_view_entries,
which projects the view's own tree.
ATOMIC_IROH_IPV4_ONLY=1 makes the owner endpoint and the client dialer
open no IPv6 socket. Under some microVM network backends (smolvm TSI, in
a fork) creating one takes the process down, so a sandbox inside such a
VM could not materialize or submit.
… or the owner's fault

A review of the remote-sandbox work turned up five defects. Three let a
remote agent do something it should not; two left a sandbox permanently
unable to work.

The CRDT id decode in export_graph_slice read the change id big-endian,
but the CRDT tables write it little-endian — the graph keys beside it are
big-endian because they are range-scanned, and the two rules sit within a
few lines of each other. id_rows drops ids it cannot resolve, so the slice
exported fine and just quietly carried no id rows for the changes that
created a file's CRDT entries. Decode through each type's own from_bytes
now, and id_rows asserts in a debug build, because a mis-decoded id is
indistinguishable from an absent one at runtime. The BE/LE rule is written
down at the encoders and in AGENTS.md.

session_is_fresh walked turns 1..=n, one read transaction per step, with n
straight off the wire — a single request could pin a runtime worker for
hours. A session's turns are one contiguous, turn-ordered run of keys, so
the last key in the range answers the same question. Both callers wanted
"has anyone claimed this id?" rather than "is this number free", which the
walk only approximated: a session with history above the asked turn read as
unused, and that is the wrong way round for a check that decides who may
write to it.

A materialized entry's path was joined onto the sandbox root with no check,
so a path that climbs out, or names the pointer or the cache, let an agent
rewrite its own trust anchor or write as the user somewhere it was never
given. The owner's forbidden_path was a reserved-name list and could not
catch `..` at all, though a submitted change lands in TREE and the next
materialize writes root.join(path).

view_tree_items fell through to TREE — the *current* view's — when it could
not open a write transaction to project another view's, so a read-only
repository answered a question about one view with another's tree. For a
remote sandbox that is a tree written to disk.

And a rejected submit was a dead end. Only a successful submit refreshed the
cache's view row, so a StaleView refusal left a sandbox that could never
land anything again, and the recovery its own error named did not exist.
The owner now sends the view as it is with the refusal, and the cache takes
it, so a retry is computed against the graph the refusal was about.

Separately, `--ttl` panicked inside the token registry while holding its
lock, and every other method took that lock with unwrap: one absurd value
poisoned the mutex and took every sandbox plus the local create/close down
with it, until the owner was restarted. The deadline is now checked
arithmetic computed before the lock is taken, and a poisoned lock is
recovered rather than propagated. A non-positive TTL is refused instead of
quietly minting a token that is already dead.

Full workspace suite: 57 binaries, 9088 tests, 0 failures. Each fix has a
test that fails without it.
atomic-canonical-wasm named the crate it wraps twice over — atomic, twice —
and it is the only wasm crate in the workspace, so the extra precision cost
a word and bought nothing. atomic-wasm says the same thing in half the
length, and the module a page imports becomes atomic_wasm.

atomic-canonical itself keeps its name: that is the crate this one wraps,
and it is a dependency, not a rename of it.

Verified: builds for the host and for wasm32-unknown-unknown (787 KB
artifact at the path build.sh expects), full workspace suite 9088 tests
across 57 binaries, 0 failures.
…found

Every remote-sandbox test so far drove one sandbox, one request at a time,
against a database nobody else was touching. A stress test with four
sandboxes — each on its own draft, since a view has one live token — all
materializing, publishing provenance and recording behind a barrier, found
two things a serial test cannot see.

The first it found immediately. FileStates and Changes opened the database
with a bare open_readonly, which fails outright when another request holds
the file; the two write handlers have been waiting on open_existing_wait all
along. So the second sandbox to ask lost, every time, with "Database already
open. Cannot acquire lock." Both read handlers wait now, for the same reason
the writes do.

The second it found by instrumenting, and it is worse. A remote cache
mints node ids from a counter seeded only by the ids the owner chose to send
it: import_ids takes fetch_max over the ids in a single import, and a cache
is only told about the files it asks for. On the four-sandbox run the owner
sent 5 ids with a maximum of 7, leaving each cache's counter at 8 while the
repository's own was well past that. The cache then invents ids for its new
content that are already real repository node ids, and the owner refuses the
change with "node N is not on this view" — the agent's work is lost, and the
sandbox cannot get past it. Parallelism is not the cause, only what makes it
reliable: it races several caches through one import so they all record from
a stale counter, and a single sandbox passes whenever the repository happens
to have no node where its cache lands. That is why nothing caught it.

Kept as an ignored test with the diagnosis, so it runs the moment the node-id
space is fixed. Fixing it is the work LOCAL_INODE_FLOOR already does for
inodes: a cache's node ids need a floor, and the owner needs to remap them
into the repository's space on insert instead of applying a submitted
change's ids verbatim. That is a change to what `insert_submitted_change`
trusts, so it is not done here.

The passing test asserts separation, which is the property a shared owner can
get wrong without erroring: every change on the draft that made it, no draft
carrying another's, dev untouched, dev's working tree not rewritten, and
each session its own with its own goal.

Full workspace suite: 57 binaries, 9089 tests, 0 failures.
… expose

Scales the stress test from four sandboxes to twenty, and adds the axis
four could not reach: a sandbox and the local repository on the *same* view.
That is the contention the submissions mutex does not cover — the mutex
serializes requests through the owner, and a local `atomic record` never
goes through it.

Twenty is stable (five consecutive runs, ~12s), and 32 and 64 also pass, so
the owner scales. The names are zero-padded because "recorded by sb-1" is a
substring of "recorded by sb-10", which made the cross-contamination check
fail against a perfectly correct owner. That was a bug in the test, not the
code, and worth writing down.

The shared-view test is a known failure, and the second half of it is data
loss. When the owner applies a sandbox's change to a view it does not update
a local working tree on that view, so a local tree that is behind reads the
view's new files as local deletions:

    $ atomic status
    On view dev
    Changes to be recorded:
        deleted:   from-sandbox.txt

from-sandbox.txt is a file the sandbox had just added and it is present on
the view — the host's tree simply does not have it. The host's next
`record -a` takes that offer, and in this test it did, on every round: by
the end the view carried both writers' changes in its log and neither
writer's files in its content. A human and an agent on one view cannot both
`record -a` without eating each other's work.

The first half of that test is a fixed bug and is why it is worth keeping:
the sandbox loses every round (three of three) and retries until it wins,
which is the stale-refusal recovery working under real contention. The log
assertions confirm both writers' changes land on the view regardless of who
won which round.

Kept as an ignored test with the diagnosis so it runs the moment the
working-tree sync is fixed.

Full workspace suite: 57 binaries, 9089 tests, 0 failures.
…king copy

Applying a sandbox's change to a view did not write it into a local working
tree sitting on that view. A tree that was left behind then saw the view's
new file as missing from disk, so `status` offered to record it as deleted:

    $ atomic status
    On view dev
    Changes to be recorded:
        deleted:   from-sandbox.txt

`from-sandbox.txt` was a file the sandbox had just added, and it was on the
view the whole time. The host's next `record -a` took that offer, and in the
test it did so on every round: by the end the view carried both writers'
changes in its log and neither writer's files in its content. A human and an
agent on one view could not both record without eating each other's work.

`insert_change` is a library function whose working-copy contract is "clean up
after deletions and moves" — it never writes new files, and every other
caller remembers to materialize afterwards. This caller did not, and a view is
not only moved by whoever is sitting on it. `insert_submitted_change` now
materializes the change's own paths, falling back to a full materialize when
it names none, and leaves a view that is not checked out alone — the user
switches to it to see its files, exactly as a pull into another view does.

The test also shows the stale-refusal recovery working under real
contention: the sandbox loses all three rounds against the local repository
and retries until it wins, which is the previous fix exercised by a genuine
race rather than a scripted refusal.

One gap left open and written down in the test rather than papered over: a
file the host has not recorded yet is not in the deferred tree journal, and
TREE is a projection of that journal, so the owner's insert reprojects TREE
and drops the host's uncommitted file before its `record -a` can pick it up.
That is a property of the projection rather than of remote sandboxes — any
`insert` into the current view with a pending local file does the same — and
it wants its own test and its own fix.

Full workspace suite: 57 binaries, 9090 tests, 0 failures.
…did not

A view shares one ambient graph with every other view, so a trunk inherited
from a parent carries branches attached by changes belonging to *other*
views — a sibling draft editing the same file adds branches to the same
trunk. The graph half of a slice filtered by the view's change set. The CRDT
half did not: export_crdt walked trunk_branches to branches to branch_leaves
to leaves and shipped every row it found, whoever created it.

So each cache received its siblings' branches. A later record reads those as
the ids of lines that already exist, names them, and the owner refuses with
`node N is not on this view` — correctly, since those nodes really do belong
to another view. The agent's work is lost with no way past it, and a sandbox
only worked while it was the sole writer of the file, which is why every
serial test passed.

Found by instrumenting the refusal rather than by reading: the refused ids
turned out to be real owner nodes with real change hashes, just not on the
submitting view. They were not a collision, and the cache was not inventing
ids — the ids it emitted for lines it already had came out of CRDT rows the
owner had wrongly sent. export_crdt now takes the view's change set and keeps
only the trunk, branch and leaf rows that set can see, matching the graph rows
beside it.

Both ignored tests now pass and neither is ignored any more:
`sandboxes_can_each_record_twice` runs twenty sandboxes through two records
each, and `a_sandbox_and_the_local_repository_can_share_a_view` has a human
and an agent racing on one view. Full workspace suite: 57 binaries, 9091
tests, 0 failures.
…eiling

CI runs `cargo test --workspace` on ubuntu, macos and windows with default
thread parallelism, and each sandbox is a process, so the stress count is now
SANDBOX_STRESS_SANDBOXES with a default of 20. An unparseable or non-positive
value falls back to the default rather than running nothing.

While measuring how far it goes, the number turned out to be worth writing
down. On one 15-core laptop, `several_remote_sandboxes_work_at_once`:

    20   passes, ~12s
    32   passes, ~17s
    64   passes, ~31s
    200  intermittent — 1 pass in 3, ~80-95s

The 200 failures are all the same thing, in the owner's `Materialize` handler:

    database owner materialize failed [materialize]: Database already open.

That is `open_existing_wait(30s)` giving up. The owner re-opens the database per
request and redb allows one holder per file, so concurrent clients queue on the
file lock and the last of them waits out the timeout while the owner streams
whole views one client at a time. A capacity limit, not a correctness bug —
nothing is corrupted, and 20 is nowhere near it. The fix is the owner holding
one `Pristine` and sharing it with every request through `open_with_pristine`,
releasing it when idle so local commands can still open the repository.

That last part is a constraint, so it is now a test.
`database_lock_exclusivity_test` runs two processes: the parent holds a
`Repository` and re-runs itself as the child, which reports via a file because
libtest's capture decides whether a child's println reaches its parent. The
answer is that the pristine's lock is exclusive across processes, and for a
read-only open as much as a read-write one:

    while one process holds the repository, a second sees: readonly-busy readwrite-busy

So the owner cannot hold the pristine for its lifetime without locking local
`status` and `record` out of the repository, which is why it opens per request
and retries, and why any handle it does keep needs an idle release. The test
asserts the exclusivity rather than printing it, and asserts the corollary that
dropping the handle frees the file — otherwise an idle release would quietly
not release anything. If redb ever moves to shared-read locking the assertion
fires and says the open-per-request retry has become redundant.
A sandbox's cache held its view flattened: shared, parentless, with the whole
ancestor union written into its own change log. `import_view_snapshot` built
that row by hand — `kind: Shared, parent: None` — and `export_view_snapshot`
fed it the union, so the cache's structure said one thing (this view is a
root with three changes) where the repository said another (this view is a
draft with none of its own, three inherited).

Everything downstream of "what is this view's own work" was then wrong.
`triage review` could not even default its target — the view had no parent —
and once given one, it counted the inherited history as the sandbox's
candidates and reported ORPHAN_CHANGE for changes the repository knew were
fine. Same view, same merkle, opposite verdicts:

    sandbox:  triage sb1 → dev  BLOCKED  3 changes  2 block
    host:     triage sb1 → dev  READY    0 changes  0 findings

The cache also invented its `dev`: the row `init` writes, empty with a null
merkle, never replaced because no ancestor ever arrived to replace it.

A skeleton now carries the view as itself — its scope, its parent, its own
change log — and each ancestor beside it, each with its own log, real state
and real parent: the draft chain and the nearest shared ancestor, exactly the
views `collect_visible_change_ids` unions. The cache reconstructs the union
from that chain the same way the repository computes it, and
`export_sandbox_skeleton` builds the rows' filter from the same chain, so
what the cache reconstructs is what the rows were filtered by. Ancestors
import first, so the parent link lands on a row that exists — and the
import's replace-by-name is what turns a cache's invented placeholder into
the real ancestor.

The cache's vault was already right — it ships as files in the view's tree
and bootstraps — so with the view structure fixed, triage in a sandbox and
triage on the repository produce identical worklists: same verdict, same
findings, same candidates, and bare `triage review` defaults its target from
the parent like it does locally.

The test compares both sides' full reports on a fresh sandbox with its own
recorded change, and asserts the candidate is the sandbox's edit rather than
the inherited base.

Full workspace suite: 59 binaries, 9104 tests, 0 failures.
@geekgonecrazy

Copy link
Copy Markdown
Contributor Author

Closing in favor of #242 (and the sandbox stack on top of it) — everything here landed elsewhere, and its central piece was superseded by a better design.

Where each part went:

  • Signing with a key held elsewhere, and rendering a view's tree without touching disk — this is Sign attestations with a key held elsewhere; render a view's tree without touching disk #242's own content now (its title is this PR's second half), including intent attest --prepare/--signed, aud-scoped tokens, and the diskless render.
  • The remote-sandbox kernel — the pristine slice, the sandbox cache, materialize, record-then-submit with the stale-view fence (slice.rs, remote_cache.rs, sandbox.rs) — was ported onto feat/libatomic-sandbox-service, stacked on Sign attestations with a key held elsewhere; render a view's tree without touching disk #242, as the SandboxService's domain: grant-checked, with the Remote sandboxes over the owner protocol, and signing with a key held elsewhere #224 fence kept. Its parity tests came with it.
  • The owner protocol is gone on purpose. The design here — the sandbox CLI dialing the host's database-owner over iroh, the pointer carrying the owner's iroh address — put transport inside atomic. The replacement keeps transport out of atomic entirely: the sandbox CLI speaks its existing daemon socket (ATOMIC_SERVICE=reactor), the host (outpost) forwards those bytes over iroh to its own in-process service layer, and the grant — minted, persisted, and checked by the host's SandboxService — is what a sandbox's token reaches. That's atomicdotdev/outpost#9, with atomic carrying no iroh anywhere.
  • The integration tests over the owner protocol were replaced by the socket-and-service coverage on the sandbox stack (remote_sandbox_client_test.rs, sandbox_service_test.rs), including the fence, forged/wrong-view token refusals, and agent-hook provenance through the grant.

So nothing here is lost — the kernel and the signing live on, and the protocol this PR was named for is the one part retired by design.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant