Conversation
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
ApprovabilityVerdict: Not approved Macroscope's review found this PR not approvable — This PR introduces cross-process state-directory locking and changes server startup, service installation, desktop restart, database recovery, and SSH process-management behavior across several production paths. Unresolved findings describe risks of terminating a reused PID and tearing down a live server after state publication fails, while the change also adds static-analysis suppressions. Not approved because:
Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more. |
fd8bd47 to
449ae8e
Compare
| CURRENT_STARTED_AT="" | ||
| case "$REMOTE_PID" in | ||
| ''|*[!0-9]*) ;; | ||
| *) CURRENT_STARTED_AT="$(LC_ALL=C ps -p "$REMOTE_PID" -o lstart= 2>/dev/null || true)" ;; |
There was a problem hiding this comment.
🟠 High src/tunnel.ts:624
The stop script can kill an unrelated process when the original server's PID is reused within the same second. ps -o lstart= only has whole-second precision, so CURRENT_STARTED_AT can equal the saved REMOTE_STARTED_AT for the replacement process; store and compare a higher-precision process start identity instead.
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @packages/ssh/src/tunnel.ts around line 624:
The stop script can kill an unrelated process when the original server's PID is reused within the same second. `ps -o lstart=` only has whole-second precision, so `CURRENT_STARTED_AT` can equal the saved `REMOTE_STARTED_AT` for the replacement process; store and compare a higher-precision process start identity instead.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 449ae8e. Configure here.
| if (typeof address === "string" || !("port" in address)) return; | ||
| const state = yield* makePersistedServerRuntimeState({ config, port: address.port }); | ||
| yield* ownership.publish(state); | ||
| }), |
There was a problem hiding this comment.
Publish failure tears down live server
Medium Severity
ownership.publish now runs after activation with no failure handling. A disk or rename error that used to only skip the discovery record now fails runtimeStateLayer and tears down an already-listening server. Pairing and T3 Connect go down with it even though the process already owns the state directory.
Reviewed by Cursor Bugbot for commit 449ae8e. Configure here.
…cked The ownership lock had two regressions for existing users. A desktop backend that lost the lock exited with code 1, so the desktop restart loop retried forever with no message. Desktop SSH stopped replacing its own remote server when the runner script changed, so an app update no longer updated the remote CLI. The refused server now exits with code 78. The desktop stops the restart loop on that code and shows a dialog that names the cause. SSH reconnect replaces a running server only when the saved PID and start time prove it is the launcher's own and the bundled runner changed. Any other server is reused as is. A publish failure after startup is a warning instead of a shutdown. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
449ae8e to
4700334
Compare


T3 Connect could start a background service while an SSH-launched server still used the same state directory. The old server's shutdown could then delete the new server's runtime record.
Acquire an OS-backed SQLite lock for the resolved state directory before server startup. Hold it through shutdown and check a unique owner ID before removing the runtime record. Service setup refuses an unmanaged takeover and gives an explicit stop-and-retry step. SSH reconnect uses the current server address without stopping a saved PID. Update backup and rollback use the same lock.
Verification
All process tests ran locally on Linux with temporary state. No live machines or relay services changed. macOS and Windows process behavior was not tested locally. Older binaries do not implement the lock, so a live legacy runtime record blocks takeover until that server stops.
The initial relay startup error remains unproven and is not attributed to this bug.
Created with GPT-6 Astra (preview) in Codex.
Takeover
Two regressions for existing users were fixed on top of the original commits, after real-process checks on Linux.
Real-process checks on the lock itself: duplicate start refused, crash releases the lock, graceful stop clears the record, legacy record with a live PID refused, reused legacy PID allowed,
node --watchrestart works, and Bun and Node respect each other's lock.Not verified: macOS and Windows process behavior, and the desktop dialog in a packaged build.
Original work by GPT-6 Astra (preview) in Codex. Takeover by Claude Fable 5.1 in Claude Code.
Note
Prevent duplicate servers for one state directory with SQLite ownership lock
acquireServerOwnershipin serverOwnership.ts backed by an exclusive SQLite transaction in the runtime-state directory; the lock is held for the server lifecycle and released on exitSERVER_EXIT_CODE_STATE_DIR_OWNED(78)backupDatabaseOnceandrestoreDatabaseBackupin serviceLauncher.ts with the same ownership lock so concurrent update/restore operations are mutually exclusiveconnect,projectdiscovery, andbootServiceinstall now check ownership before proceeding; the desktop backend in DesktopBackendManager.ts stops permanently instead of restarting when the backend exits with the ownership statusisProcessAlivein serverRuntimeState.ts returns false for non-positive or non-integer PIDs;persistServerRuntimeStateandclearPersistedServerRuntimeStateare removed in favor of the ownership modulePersistedServerRuntimeStategains an optionalownerIdfield andServerRuntimeStateErrorno longer covers persistence/clear operations — any out-of-tree callers of the removed functions or the old error cases will breakMacroscope summarized 4700334.
Note
High Risk
Changes core server lifecycle, discovery records, service install, and SSH cleanup across a shared state directory; mis-ownership could still allow duplicate servers or block legitimate restarts until legacy servers stop.
Overview
Introduces SQLite-backed state-directory ownership so only one T3 Code server can use a given home at a time. Startup acquires an exclusive lock (
server-owner.sqliteon the realpath-resolved userdata dir), publishesserver-runtime.jsonwith anownerId, and only removes the discovery file on shutdown when that ID still matches—fixing races where Connect/service install could run a second server and an old process could delete the newer runtime record.CLI and service behavior now refuse silent takeover:
t3 connectsaves auth but skips background install if a server still holds the directory; boot service install callsrequireServerStoppedbefore first install and again after stopping the unit. Project commands treat dead PIDs as offline and no longer clear the runtime file when a live-server request fails.SSH remote launch prefers the current
server-runtime.jsonendpoint without killing saved PIDs, records process start time for managed launches, and only stops or blocks reconnect when start times match. The service launcher takes the same lock during SQLite backup/restore so updates do not race a foreground server.Docs and broad integration tests cover concurrent starts, legacy records, and the original incident scenario.
Reviewed by Cursor Bugbot for commit 449ae8e. Bugbot is set up for automated code reviews on this repo. Configure here.