Skip to content

fix(orchestrate): tell the orchestrator why a revived child failed to start - #604

Merged
Tryanks merged 1 commit into
mainfrom
fix/child-start-failure-callback
Oct 6, 2026
Merged

Tryanks merged 1 commit into
mainfrom
fix/child-start-failure-callback

Conversation

@Tryanks

@Tryanks Tryanks commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Summary

An orchestrator send to a long-idle child whose worktree had been removed:

 send(child)                          → {"delivery":"queued"}
   start_session(cwd = removed worktree)
-    spawn → ENOENT → "spawning `…/claude`: No such file or directory"
+    "working directory `…` no longer exists"
   ProviderStartFailed folds into the already-reported turn
   deliver_child_callback(Failed)
-    same turn count → deduplicated, parent hears nothing (app kept running)
-    after a restart: "<thread> failed." + the old turn's final message
+    "[orchestrate] thread X (\"title\") failed to start: <error>"
 status(child)
-  state: completed, last_output_tail: old message
+  state: failed, start_error: <error>
  • crates/agent/src/lib.rs: start_session rejects a missing cwd before any provider spawns, so all providers report it the same way.
  • crates/runtime/src/app/orchestrate.rs: trailing_start_error (the timeline's last entry is a ProviderStartError) drives the one-line callback, bypasses the per-turn dedupe, and sets the state that status reports.

Evidence

Found in real use: child 52348392 was revived after 3 days. Its cwd ~/.tcode/worktrees/native-grok was gone, the error blamed the claude binary (which still existed), and the parent's callback carried the 3-day-old final message.

  • Before: send_to_child_whose_cwd_was_removed_reports_the_start_failure fails. run_until times out after 5s because the parent never receives a callback.
    After: the test passes. The parent receives exactly [orchestrate] thread child ("…") failed to start: failed to spawn provider process: working directory … no longer exists, and status reports failed with start_error.
  • cargo fmt --all --check, cargo clippy --workspace --all-targets --locked -- -D warnings and cargo nextest run --workspace --locked (937 passed) were run locally on macOS. Machete and the mobile/Web checks are left to CI.

Merge Danger

Door: two-way

Blast Radius: narrow

Two things change. Every provider start now stats its cwd first. The orchestrate callback and status text change for a child whose last start failed; ordinary completed/failed callbacks are unchanged. No wire or persisted-format change.

… start

An orchestrator returning to a long-idle child whose worktree had since
been removed got a misleading or missing answer. The spawn into the
missing directory failed with the OS's "not found", which the provider
reported as its binary being missing. The failure folded into the turn
that had already been reported, so its callback was deduplicated away
while the app kept running. After a restart it did arrive, but it carried
that old turn's final message instead of the error, and `status` still
said `completed`.

`start_session` now refuses a missing working directory by name. A child
whose last entry is a failed start reports `failed to start: <error>` in
one line, is not deduplicated against the earlier turn, and shows as
`failed` with `start_error` in `status`.
@Tryanks
Tryanks merged commit bcbc71a into main Oct 6, 2026
7 checks passed
@Tryanks
Tryanks deleted the fix/child-start-failure-callback branch October 6, 2026 11:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant