Skip to content

Engine: settle live runs together on shutdown - #792

Open
RemingtonWilcox wants to merge 1 commit into
zeronsh:mainfrom
RemingtonWilcox:quit-interrupts-together
Open

RemingtonWilcox wants to merge 1 commit into
zeronsh:mainfrom
RemingtonWilcox:quit-interrupts-together

Conversation

@RemingtonWilcox

@RemingtonWilcox RemingtonWilcox commented Oct 4, 2026 •

Copy link
Copy Markdown

Summary

  • Sessions::shutdown interrupted live runs one after another, and each interrupt waits for its run to settle. A run whose harness doesn't wind down settles at the engine's 3 s interrupt deadline, so stopping an engine with N live turns took about 3 s per turn.
  • The interrupts now run concurrently (join_all, as DocHost::shutdown_workers already does), so shutdown costs one bounded settle in total however many turns are live. Each run still settles the same way: streaming entries are stamped aborted and the journal is closed.
  • Where it shows: zeron headless stopping (Ctrl-C, StopEngine, zeron daemon stop), and the desktop's runtime switch, which gives the daemon 10 s to stop. With 4 or more live turns that switch used to time out with "Could not stop the remote engine".
  • No protocol or behaviour change beyond timing.

Worth a close look

  • This does not fix "the window closed but Zeron kept running" on Windows, and I couldn't reproduce that. On current main, closing the window with live turns exits in about 0.3 s. GPUI gives on_app_quit handlers 200 ms (SHUTDOWN_TIMEOUT), logs timed out waiting on app_will_quit, and drops the task. That drop aborts the in-process engine shutdown through gpui_tokio. Process exit then drops the Tokio runtime. I measured this with 3 live mock turns, and again with 3 turns whose agent CLI was a stand-in process that never writes or exits. In both cases the process exited and no agent processes were left behind. So the headed quit never waits for live turns: it doesn't hang, but it also doesn't finish the graceful drain (aborted stamps, final snapshot flush). Recovery on the next boot covers that today. I left it alone because it's a separate design question.
  • I read every await in EngineCore::shutdown and InProcessEngine::shutdown and found none without a bound that the quit path could block on, so this PR adds no extra timeout.
  • One known way to hang that I didn't trigger: dropping the Tokio runtime waits without a limit for blocking-pool work. On Windows, agent stdout/stderr pipes are tokio::fs::File, so their reads run on the blocking pool. If something kept an agent child alive past runtime shutdown, the process could linger. With the stand-in agent, the run tasks' Child drop terminated the job and nothing lingered.
  • The new test spends about 3 s of real time (the engine's interrupt deadline).

Test plan

  • New e2e::shutdown_settles_live_runs_together: 3 live runs whose harness ignores the interrupt; EngineCore::shutdown must finish in under 6 s and stamp all three aborted. Passes in about 3.3 s. Without the fix it fails at about 9.4 s (checked 4 times).
  • cargo test -p zeron-engine --test e2e on Windows: everything else passes, except generated_image_is_materialized_before_publication_and_survives_reopen. That test also fails on unmodified main here. wrong_id_respond_is_rejected_and_correct_answer_still_resumes, harness_emitted_input_twin_is_dropped_and_answer_resumes and interrupt_stamps_streaming_entry_aborted failed now and then in full parallel runs, both with and without this change, and pass when run alone.
  • cargo test -p zeron-engine --lib sessions
  • Manual, Windows debug build with a throwaway data dir, mock harness (ZERON_MOCK_DELAY_MS=1000, ZERON_MOCK_REPEAT=50), live turns started over IPC:
    • zeron headless + StopEngine, 3 turns: 9.65 s before, 3.63 s after
    • same, 6 turns: 18.6 s before, 3.6 s after
    • same, 3 turns with a stand-in agent CLI that never answers: 6.8 s before, 2.6 s after
    • headed, closing the window with 3 live turns (mock or stand-in agent): about 0.3 s, before and after
  • Not checked on macOS or Linux.
  • Not checked with a real Claude/Codex turn.
  • Linux UI and core suites on the fork's CI (ui-tests workflow): https://github.com/RemingtonWilcox/zeron/actions/runs/37240703827

Screenshots

None: no visible change.

🤖 Generated with Claude Code

Engine shutdown interrupted live runs one at a time, and each interrupt
waits for its run to settle, so stopping an engine with N live turns
took about 3 s per turn. Interrupt them concurrently so it costs one
bounded settle in total.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant