Skip to content

/arch-save: packaged save→clear→re-init cycle for architect context refresh (counterpart to /arch-init) #1307

Description

@waleedkadous

Proposal (from a production workspace's architect, live-tested there)

Long architect sessions accumulate stale context. The refresh recipe exists today as three manual steps, and /arch-init's skill doc already documents the loop as prose (save at a resumable boundary → suggest /clear → human clears → /arch-init recovers). Package it as a command — /arch-save (or afx reset <architect> --state) that:

  1. Writes the dated resume block into codev/state/<name>.md — current lanes, pending confirmations, next actions, and (live-run finding) an explicit list of session-bound monitors/watchers that die at the clear and need re-arming — everything /arch-init needs to resume cold. Write happens strictly BEFORE any clear.
  2. Clears the architect's own session.
  3. Re-injects /arch-init <name> into the fresh session so it re-adopts identity and resumes automatically.

Design notes

  • Reuse Builder context reset should be a first-class flow: save-state → /clear → re-orient #1273's machinery, not a detached scheduler. The proposal's original leg-3 design (detached process sleeping ~45s to send /arch-init, because the sender's session dies at the clear) predates PR [Spec 1273] Builder context reset: save-state → /clear → re-orient #1305: afx reset already implements Tower-owned interrupt → /clear via the raw/escape channel → post-clear confirmation → re-orientation injection, for builders. /arch-save is the architect flavor: same Tower-side sequencing, but the re-orientation payload is /arch-init <name> (state-file-based recovery) instead of spawn-prompt reconstruction. Tower survives the clear; no orphan scheduler process needed.
  • The human-keystroke invariant must survive packaging. /arch-init's save discipline deliberately keeps the irreversible step behind a human decision: an agent must never unilaterally destroy its own context. /arch-save doesn't break that — it relocates the human decision from "press /clear" to "invoke /arch-save" — but the command doc must say so: architects run it on the owner's direction or the owner runs it; an architect must not invoke it autonomously mid-task on its own judgment. (Standard override-carveout framing: "don't autonomously X," not "X is forbidden.")
  • Refuse mid-task saves. The skill doc's rule — save only at a resumable boundary, never mid-task — is the part a command CAN'T verify. At minimum the command should require a --boundary-style acknowledgment; the quality of the resume block stays on the architect.
  • Gate on the Builder context reset should be a first-class flow: save-state → /clear → re-orient #1273 verify e2e. The underlying reset headline path (does /clear actually take effect via the raw channel; what a real clear emits) is still pending its live end-to-end run. /arch-save inherits that dependency.

Evidence

The proposing workspace is live-running exactly this cycle manually (their cost architect, state file at v67) and has offered the v67 state block as a template for the resume-block format. Their run surfaced the monitor-re-arm requirement and the write-before-clear ordering as real failure modes, not theory.

Activity

  1. waleedkadous commented on Jul 31, 2026

    @waleedkadous
    ContributorAuthor

    Resume-template candidate: the v67 state block (sanitized)

    As requested — the state block that drove the live save→/clear→/arch-init cycle on 2026-07-31. One substitution: this repo is public, and the live block carries internal program content (spend figures, un-released experiment results, verbatim stakeholder quotes). Below is the same block with content genericized; every structural element is preserved verbatim from the live artifact. The un-sanitized original lives in our project at codev/state/cost.md (v67 header) for anyone with repo access.

    What the live run validated, in template order:

    1. Intent stamp first — the block opens by declaring the /clear is deliberate and owner-directed, so the resuming instance never treats it as a crash.
    2. Monitor re-arm list — session-bound monitors die with the clear; the block enumerates exactly what to re-arm (watch targets + cadence + alert pattern). This fired correctly on resume.
    3. DONE-with-receipts — everything completed pre-clear is listed with its verification evidence (merge SHA verified on origin, push-verified branch SHA), so the resuming instance doesn't re-do or re-verify from scratch.
    4. ACTIVE lanes with brief pointers — each in-flight workstream names its builder id + the on-disk brief file, so no instruction content lives only in the lost context.
    5. Latest results block — the current decision-relevant numbers, so the first post-resume decision needs no archaeology.
    6. Queued-with-ordering — deferred work with its explicit sequencing dependency ("waits for X verdict").
    7. Authorization envelope — what spend/authority survives the clear vs expired with the completed work.
    # <lane> architect — state (vNN, <date> ~HH:MM UTC — <milestone>, DELIBERATE /clear cycle)
    # ⭐ THIS /clear IS INTENTIONAL (owner-directed context refresh). On re-init: normal
    # /arch-init flow, then:
    # 1. RE-ARM MONITOR (old ones died with the clear): <cadence> HB on <builder-ids>
    #    liveness (pattern in vNN-2 WATCH block below).
    # 2. DONE pre-clear: <PR> MERGED (<sha>, verified on origin/<branch>; root synced).
    #    <branch> PUSH-VERIFIED (<sha> local==origin) and worktree RELEASED. <worktree>
    #    still HELD (<who> reads its receipts; told to stand by).
    # 3. ACTIVE LANES (owner-directed wave, spawned ~HH:MM with brief files in <dir>):
    #    <builder-a> = <workstream> (<brief-file> + <addendum summary>);
    #    <builder-b> = <workstream> (<brief-file>; <standing rule>);
    #    <builder-c> = <workstream> (<brief-file>; <standing rule first>).
    # 4. <MILESTONE> FINAL: <headline numbers>, <verdict>. Full close-out in vNN-1
    #    block below. Spend ~$X. Owner reaction: <quote>.
    # 5. MERGED TODAY post-report: <PRs + shas + issue closures>.
    # 6. QUEUED (owner-approved sequencing): <item> — WAITS for <verdict>; <item> —
    #    <when>; <item> still open with owner.
    # 7. Envelope: <standing authorization>; <expired authorization> expired with the
    #    completed work.
    #
    # (vNN-1 header follows)
    

    Two live-run findings already cited in the issue body, confirmed from this side:

    • Write-before-clear ordering is the whole game — the block was written and committed before the keystroke; the resuming instance had zero reconstruction to do.
    • Session-bound monitors must be listed for re-arm, not described as "in place" — on resume the re-armed monitor's first alert was a false positive (transient afx status race); the replacement needed a 2-consecutive-miss debounce. Worth encoding in the /arch-save flavor: a freshly re-armed monitor should self-test its predicate once before its alerts are trusted.

    On amendment (1): agreed, our detached-scheduler leg is obsolete if afx reset owns interrupt→clear→re-orientation — inheriting reset's pending live-e2e gate is the right coupling. On amendment (2): the human-keystroke invariant is the part we'd insist on; autonomous self-invocation ban with a standard override carveout matches how the live run was actually operated.

  2. waleedkadous commented on Jul 31, 2026

    @waleedkadous
    ContributorAuthor

    Correction to the monitor finding (from the same live run) — changes the spec

    Earlier evidence on this issue said session-bound monitors die at the clear and need re-arming. The live run disproved half of that: monitors are session-bound, not context-bound — a watcher armed in the pre-clear context SURVIVED /clear and fired a stale false alert 8 minutes into the fresh context, against a target that had been deliberately decommissioned before the clear. (The earlier "monitors die" observation came from a night where the sessions themselves died, which conflated the two lifecycles.) Process-level orphan checks (pgrep) can't see these — they're harness background tasks, not shell processes.

    Spec change for /arch-save's post-clear checklist, order is load-bearing:

    1. Enumerate + STOP stale monitors first — everything armed by the pre-clear context is keyed to a world the fresh context can't verify, and a surviving watcher's alert is indistinguishable from a fresh one.
    2. Then re-arm the monitors listed in the state block, with the previously-noted first-check debounce/self-test.

    Equivalently: the state block's monitor list serves two purposes — a kill-list for the transition and a re-arm list for the resumed instance — and /arch-save should treat it as both.

  3. waleedkadous commented on Sep 22, 2026

    @waleedkadous
    ContributorAuthor

    Shipped: the /arch-save cycle landed via Spec 1307 (commits b4c5d08, 45f946e) and the next-task extension via #1709 / PR #1710 (merge 69af908). Closing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/towerArea: Tower server / agent farm CLI

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions