Repository navigation
afx workspace recover: revive builders after machine reboot/crash #829
Copy link
Copy link
Closed
Labels
area/towerArea: Tower server / agent farm CLIArea: Tower server / agent farm CLI
Description
Activity
- addedtype:featureNet-new capabilityNet-new capabilityarea/towerArea: Tower server / agent farm CLIArea: Tower server / agent farm CLIand removedtype:featureNet-new capabilityNet-new capability
on May 23, 2026 Status update — implementation landed on
feat/829-workspace-recover, draft PR #833.What this PR ships (closes #829):
afx workspace recovercommand — enumerates porch projects, identifies dead builders, respawns viaafx spawn --resume. Dry-run by default.- Eligibility predicate handles terminal/unsupported/worktree-missing/alive/stale.
- Conversation resume via on-disk jsonl discovery for all supported builder protocols (SPIR/ASPIR/AIR/PIR/bugfix) — see Builder conversation resume: restore Claude session on recovery via jsonl discovery #831 for that work specifically.
- Restricted to revivable protocols (
spir,aspir,air,pir,bugfix);experiment/maintain/task/legacyspiderskipped withunsupported_protocolreason.
Three rounds of cmap-3 review:
- Round 1 caught four blockers — fixed in
69d65d37. - Round 2 surfaced the session-row-existence predicate bug (real blocker — was breaking the primary use case against shannon) plus conversation resume scope expansion; fixed across
c14fd0d1,c6c0dc98,386353d4,b96dcacc. - Round 3 (just fired) reviews the round-2 fixes themselves:
discoverResumeSessionhelper extraction so PIR/bugfix recovery also picks up prior conversations, and a defensive guard so main architect resume can't steal a sibling's session in multi-architect workspaces.
Live-tested against shannon (345-project workspace) — 9 in-flight PIR/SPIR builders correctly surfaced as
revive; rest classified across terminal/unsupported/missing/stale.Related issues:
- Workspace recover: revive Tower-managed architects after machine reboot #830 — architect session revival (main shipped here; siblings → Multi-architect conversation resume: disambiguate via per-architect session UUID #832)
- Builder conversation resume: restore Claude session on recovery via jsonl discovery #831 — builder conversation resume mechanism
- Multi-architect conversation resume: disambiguate via per-architect session UUID #832 — multi-architect per-architect UUID storage (Waleed)
- added a commit that references this issue
on May 24, 2026 - added 7 commits that reference this issue
on May 28, 2026
Metadata
Metadata
Assignees
Labels
area/towerArea: Tower server / agent farm CLIArea: Tower server / agent farm CLI
Problem
After a machine reboot or crash, all builder processes die — shellper daemons (detached, in-memory) and Tower together. Disk-persistent state survives: worktrees in
.builders/<id>/, porchstatus.yaml, SQLiteterminal_sessions, spec/plan/review files. But there is no command to walk that state and respawn the builders that should still be running.The user typically discovers a half-dozen "dead" builders the morning after a reboot and has to
afx spawn <id> --resumeeach one by hand. Most builders are sitting idle at approval gates when this happens, so a fresh respawn (which lands them right back at the gate viastatus.yaml) is functionally equivalent to true session resume.Proposal
Add an explicit
afx workspace recovercommand (no auto-trigger on Tower startup) that:codev/projects/*/status.yaml(usingfindStatusPath()to handle spec-653 multi-PR layouts under.builders/<id>/codev/projects/).--applyto actually respawn.afx spawn <id> --resumecodepath.Revival predicate
A project is eligible when ALL of:
phase ∉ {verified, complete}— not in a terminal porch state.terminal_sessionsrow exists for this project (it had a builder previously)..builders/<id>/still exists on disk.status.yaml.updated_atis within the last 7 days (configurable via--max-age <days>; bypass entirely with--include-stale).Notes:
prgate are revived (consistent with all other gate-idle builders; the resource cost is negligible and the human may want to dispatch the builder to address review feedback without manual respawn).status.yamlwas never written (builder crashed before first phase transition) won't show up; this is the conversation-resume gap and is out of scope.CLI
Out of scope (separate issue)
~/.claude/projects/<encoded-cwd>/<uuid>.jsonlfile). Requires capturing the Claude session UUID at spawn time, storing it interminal_sessions, and passing--resume <uuid>to the relaunchedclaudeprocess. Worth it for mid-phase recovery but materially larger scope; file separately if/when mid-phase crashes become a real pain point.Implementation notes
findStatusPath()(state.ts:285–303) for status.yaml lookup.terminal_sessionsrow → PID alive + socket connectable. Reconciliation logic intower-terminals.ts:485–711already detects dead sockets; recovery extends that path with respawn instead of just cleanup.afx spawn <id> --resumerather than building a parallel codepath.