Repository navigation
[Bug]: Stopping a duplicate backend cancels runs on the surviving server sharing the same T3 home #17003
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Thanks for the careful write-up, @seankoji. I haven't reproduced this in a running app, but both effects you saw match the code on current
main(3143335). They are the shutdown side of #6097. Shutdown and startup cleanup both assume one backend owns the T3 home, so a second backend's cleanup reaches into the first one's live runs.1. A graceful shutdown cancels every live run in the shared database
- On shutdown, the server runs
prepareForShutdown, stops its own provider sessions, then callsreconcile("shutdown")(serverRuntimeStartup.ts#L435-L461). reconcileloops over every thread returned bygetRecoveryThreadIds("runtime")(ProviderRuntimeRecoveryService.ts#L723-L759). That query selects any run inpreparing/starting/running/waiting, with no owner filter (ProjectionStore.ts#L3513-L3524).- For each of those runs it writes
status: "cancelled"(#L319-L327), with the detail "Cancelled because the server shut down before the provider work completed" (#L238). It does the same to attempts, nodes, provider turns and turn items. It also marks every non-stopped provider sessionstopped(#L617-L628). The comment at #L586 says why: "All provider processes are gone on startup/shutdown". - The effect outbox does the same.
reconcileAfterProcessLosscancels every pending or running process-bound effect regardless oflease_owner(EffectOutbox.ts#L483-L497). - Runs have no per-process owner (pid, instance id or lease). The outbox's per-worker lease is the only owner-like field I found, and the cleanup above ignores it.
- Startup does the same thing.
recoverisreconcile("startup")(#L817-L819), so a second backend starting on a home with a live backend would also cancel that backend's runs.
Why the tools refused. Provider-scoped MCP calls require the calling thread to have an active run (
preparing/starting/running/waiting) on the calling provider instance. Otherwise they fail with exactlyThe calling provider no longer owns an active thread run.(threadAccess.ts#L116-L133, OrchestratorMcpService.ts#L1037-L1049, ThreadManagementService.ts#L349-L356). After the other backend cancelled the run row, the surviving backend's live Codex turn had no active run, so every check failed.Possibly related to step 5 (not verified). When "continue threads after server update" is on,
prepareForShutdownalso queues aprovider-runtime.continueeffect in the shared outbox for each thread that looks resumable, including threads driven by the other backend (#L777-L803). I haven't checked whether the surviving backend's effect worker picked those up and caused the failed continuation attempts you saw.2. Shutdown deletes the discovery file without checking who wrote it
server-runtime.jsonrecords the writer'spid(serverRuntimeState.ts#L56-L70), and each server writes it on activation.- The release step calls
clearPersistedServerRuntimeState(config.serverRuntimeStatePath)without reading it first (server.ts#L721-L754). That function is a plainfs.remove(path, { force: true })(serverRuntimeState.ts#L90-L113). There's no pid or owner comparison, so any backend that shuts down removes the record, even when it was written by a different backend that is still running. - That also turns off the one existing guard. The
t3/t3 startpreflight refuses to start over a live server by reading this same file (cli/config.ts#L333-L342). Its own comment calls it advisory, not a lifetime lock. It only runs inwebmode fort3/t3 start, not fort3 serve(cli/server.ts#L60-L70) or the desktop backend. Once the record is deleted, even that check passes.
Confirmed vs. not
- Confirmed in code: shutdown and startup reconciliation cancel every non-terminal run, session and process-bound effect in the shared database, with no ownership check. The discovery file is deleted without checking who owns it. The MCP error fires whenever the calling thread has no active run.
- Not reproduced: the incident itself, its timing, the mixed-build factor, and whether the queued continuation effects caused step 5.
Fix options (for maintainers to choose)
- Single owner per T3 home. This is the root cause. Closed PRs #6098 and #9652 weren't rejected on principle. fix(desktop): prevent shared database ownership #6098 was closed because a probe-based preflight isn't atomic ("the duplicate-owner bug is valid, but the fix needs atomic ownership"). fix(server): prevent duplicate servers for one state directory #9652 was closed only because it conflicted with the V2 rewrite. Open #16102 takes exclusive ownership before persistence opens and adds an
ownerIdto the runtime record. If it lands, a second backend can't start, which should make both effects above unreachable. It's worth retesting against this scenario once it does. - Defense in depth, if wanted: only remove
server-runtime.jsonwhen itspid/owner id matches this process, which fix(server): enforce one database owner and guard update rollback #16102 appears to cover. Scoping shutdown reconciliation to runs and sessions this process actually started would need a per-process owner on runs, which doesn't exist today.
Related
- #6097: the same root cause, multiple backends on one database. This report adds the effect on the other backend's live runs.
- #5749: SSH reconnect starting a competing server on the same home.
- #14189: the Codex
already has an active writererror, which #6097 comments also tie to duplicate backends. - #15357: the opposite case, where startup recovery cancels runs whose providers outlived a killed backend. The cause is the same: recovery assumes every provider process is gone.
- On shutdown, the server runs
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 7, 2026
Summary
On Linux Mint, multiple T3 backends were running against the same default
~/.t3directory and environment ID. During cleanup, gracefully stopping the older backends coincided with active runs on the surviving managed server being marked cancelled, even though that server and its provider processes kept running.The native Codex turn continued executing, but its durable T3 run state became failed after an automatic continuation attempt. Provider-scoped browser tools then refused calls with:
Related: #6097 covers multiple backends sharing one database. This report adds the observed shutdown effect on another backend's active runs. #15357 concerns providers surviving a killed owning backend; in this incident the active provider's owning backend remained alive.
Environment
t3 serve: nightly build ending.2761.0.0.46-nightly.20261007.2787.Observed sequence
The cleanup itself triggered the incident. Checking only the stopped backends' child processes was insufficient to establish that their shutdown would leave other runs unaffected.
Expected behavior
Only one backend should be allowed to own a T3 home. If duplicate owners exist, one backend's shutdown must not cancel runs or remove the runtime descriptor belonging to another live backend.
Investigation / reproduction limits
This is one observed live incident, supported by process inventory, database state, provider activity, and the tool ownership error. It has not been reproduced in an isolated test, and no specific source-code cause is claimed. The mixed builds are a relevant variable.
A candidate isolated reproduction is to run two backends with one disposable T3 home, start a controlled provider turn on backend B, then gracefully terminate idle backend A. Check whether B's run stays active, its provider calls remain authorized, and its runtime descriptor remains valid. Please do not reproduce against a live home with important work.
Recovery / workaround
The desktop was switched to client-only mode (
localEnvironmentEnabled=false), leaving one managed backend. After the native audit turn and active runs had settled, the managed service was restarted.Fresh checks confirmed one backend and one relay client, matching local and public relay discovery responses, SQLite
quick_checkpassing, working browser automation, and newly running provider turns. Interrupted tasks were not automatically reissued or confirmed completed.Raw databases, credentials, account identifiers, and private relay URLs are omitted.
Investigated and drafted with Codex running in T3 Code.