Skip to content

[Bug]: Stopping a duplicate backend cancels runs on the surviving server sharing the same T3 home #17003

Description

@seankoji

Summary

On Linux Mint, multiple T3 backends were running against the same default ~/.t3 directory and environment ID. During cleanup, gracefully stopping the older backends coincided with active runs on the surviving managed server being marked cancelled, even though that server and its provider processes kept running.

The native Codex turn continued executing, but its durable T3 run state became failed after an automatic continuation attempt. Provider-scoped browser tools then refused calls with:

The calling provider no longer owns an active thread run.

Related: #6097 covers multiple backends sharing one database. This report adds the observed shutdown effect on another backend's active runs. #15357 concerns providers surviving a killed owning backend; in this incident the active provider's owning backend remained alive.

Environment

  • Linux Mint, same OS user and default T3 home.
  • Desktop AppImage backend and standalone t3 serve: nightly build ending .2761.
  • Surviving systemd-managed backend: 0.0.46-nightly.20261007.2787.
  • Orchestration V2; the directly observed continuing turn used Codex.
  • Three backends shared the same T3 home and environment ID, on ports 3773, 3774, and 37377. Each also had a Cloudflare relay client.

Observed sequence

  1. The managed backend had active provider turns.
  2. An audit found the two older backends sharing its home. Their process trees had no coding-agent children, so they were treated as idle.
  3. The standalone backend received SIGTERM and the older desktop was asked to terminate through its parent process. The managed backend was left running.
  4. Around 23:10:49-50 UTC on 7 Oct 2026 (10:10:49-50 AEDT, Thu 8 Oct), active runs in the shared V2 database were marked cancelled while the surviving service's provider processes continued.
  5. Automatic continuation attempts around 23:10:52-57 UTC on 7 Oct 2026 (10:10:52-57 AEDT, Thu 8 Oct) failed. The audit's native Codex turn continued, but T3 tool ownership checks rejected it.
  6. An older backend also removed the shared runtime discovery file despite the managed backend remaining alive.

The cleanup itself triggered the incident. Checking only the stopped backends' child processes was insufficient to establish that their shutdown would leave other runs unaffected.

Expected behavior

Only one backend should be allowed to own a T3 home. If duplicate owners exist, one backend's shutdown must not cancel runs or remove the runtime descriptor belonging to another live backend.

Investigation / reproduction limits

This is one observed live incident, supported by process inventory, database state, provider activity, and the tool ownership error. It has not been reproduced in an isolated test, and no specific source-code cause is claimed. The mixed builds are a relevant variable.

A candidate isolated reproduction is to run two backends with one disposable T3 home, start a controlled provider turn on backend B, then gracefully terminate idle backend A. Check whether B's run stays active, its provider calls remain authorized, and its runtime descriptor remains valid. Please do not reproduce against a live home with important work.

Recovery / workaround

The desktop was switched to client-only mode (localEnvironmentEnabled=false), leaving one managed backend. After the native audit turn and active runs had settled, the managed service was restarted.

Fresh checks confirmed one backend and one relay client, matching local and public relay discovery responses, SQLite quick_check passing, working browser automation, and newly running provider turns. Interrupted tasks were not automatically reissued or confirmed completed.

Raw databases, credentials, account identifiers, and private relay URLs are omitted.

Investigated and drafted with Codex running in T3 Code.

Activity

  1. juliusmarminge commented on Oct 7, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the careful write-up, @seankoji. I haven't reproduced this in a running app, but both effects you saw match the code on current main (3143335). They are the shutdown side of #6097. Shutdown and startup cleanup both assume one backend owns the T3 home, so a second backend's cleanup reaches into the first one's live runs.

    1. A graceful shutdown cancels every live run in the shared database

    • On shutdown, the server runs prepareForShutdown, stops its own provider sessions, then calls reconcile("shutdown") (serverRuntimeStartup.ts#L435-L461).
    • reconcile loops over every thread returned by getRecoveryThreadIds("runtime") (ProviderRuntimeRecoveryService.ts#L723-L759). That query selects any run in preparing/starting/running/waiting, with no owner filter (ProjectionStore.ts#L3513-L3524).
    • For each of those runs it writes status: "cancelled" (#L319-L327), with the detail "Cancelled because the server shut down before the provider work completed" (#L238). It does the same to attempts, nodes, provider turns and turn items. It also marks every non-stopped provider session stopped (#L617-L628). The comment at #L586 says why: "All provider processes are gone on startup/shutdown".
    • The effect outbox does the same. reconcileAfterProcessLoss cancels every pending or running process-bound effect regardless of lease_owner (EffectOutbox.ts#L483-L497).
    • Runs have no per-process owner (pid, instance id or lease). The outbox's per-worker lease is the only owner-like field I found, and the cleanup above ignores it.
    • Startup does the same thing. recover is reconcile("startup") (#L817-L819), so a second backend starting on a home with a live backend would also cancel that backend's runs.

    Why the tools refused. Provider-scoped MCP calls require the calling thread to have an active run (preparing/starting/running/waiting) on the calling provider instance. Otherwise they fail with exactly The calling provider no longer owns an active thread run. (threadAccess.ts#L116-L133, OrchestratorMcpService.ts#L1037-L1049, ThreadManagementService.ts#L349-L356). After the other backend cancelled the run row, the surviving backend's live Codex turn had no active run, so every check failed.

    Possibly related to step 5 (not verified). When "continue threads after server update" is on, prepareForShutdown also queues a provider-runtime.continue effect in the shared outbox for each thread that looks resumable, including threads driven by the other backend (#L777-L803). I haven't checked whether the surviving backend's effect worker picked those up and caused the failed continuation attempts you saw.

    2. Shutdown deletes the discovery file without checking who wrote it

    • server-runtime.json records the writer's pid (serverRuntimeState.ts#L56-L70), and each server writes it on activation.
    • The release step calls clearPersistedServerRuntimeState(config.serverRuntimeStatePath) without reading it first (server.ts#L721-L754). That function is a plain fs.remove(path, { force: true }) (serverRuntimeState.ts#L90-L113). There's no pid or owner comparison, so any backend that shuts down removes the record, even when it was written by a different backend that is still running.
    • That also turns off the one existing guard. The t3 / t3 start preflight refuses to start over a live server by reading this same file (cli/config.ts#L333-L342). Its own comment calls it advisory, not a lifetime lock. It only runs in web mode for t3 / t3 start, not for t3 serve (cli/server.ts#L60-L70) or the desktop backend. Once the record is deleted, even that check passes.

    Confirmed vs. not

    • Confirmed in code: shutdown and startup reconciliation cancel every non-terminal run, session and process-bound effect in the shared database, with no ownership check. The discovery file is deleted without checking who owns it. The MCP error fires whenever the calling thread has no active run.
    • Not reproduced: the incident itself, its timing, the mixed-build factor, and whether the queued continuation effects caused step 5.

    Fix options (for maintainers to choose)

    1. Single owner per T3 home. This is the root cause. Closed PRs #6098 and #9652 weren't rejected on principle. fix(desktop): prevent shared database ownership #6098 was closed because a probe-based preflight isn't atomic ("the duplicate-owner bug is valid, but the fix needs atomic ownership"). fix(server): prevent duplicate servers for one state directory #9652 was closed only because it conflicted with the V2 rewrite. Open #16102 takes exclusive ownership before persistence opens and adds an ownerId to the runtime record. If it lands, a second backend can't start, which should make both effects above unreachable. It's worth retesting against this scenario once it does.
    2. Defense in depth, if wanted: only remove server-runtime.json when its pid/owner id matches this process, which fix(server): enforce one database owner and guard update rollback #16102 appears to cover. Scoping shutdown reconciliation to runs and sessions this process actually started would need a per-process owner on runs, which doesn't exist today.

    Related

    • #6097: the same root cause, multiple backends on one database. This report adds the effect on the other backend's live runs.
    • #5749: SSH reconnect starting a competing server on the same home.
    • #14189: the Codex already has an active writer error, which #6097 comments also tie to duplicate backends.
    • #15357: the opposite case, where startup recovery cancels runs whose providers outlived a killed backend. The cause is the same: recovery assumes every provider process is gone.
  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions