Repository navigation
[Bug]: SwiftUI: an unresponsive passive environment delays cold start of the active one #14699
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Thanks for the cold-launch timings and the clear comparison between refused and silent connections, @saphid! This is a real bug. I confirmed it on
t3code/rebuild-mobile-app-swiftatbb01adbf90, the branch tip you named (open PR #5178). It only affects the experimental SwiftUI client. The React Native app we ship doesn't use this startup path.To answer your question: cold start should show the active environment as soon as its own shell is ready, and passive environments should fill in afterwards. Waiting for every enabled environment before leaving "Connecting to T3 Code" isn't the intended behavior. That matches foreground resume, which already says other saved computers must not delay recovery.
What happens today
FeatureRootModel.isLoadingstarts out true, andFeatureRootViewshows "Connecting to T3 Code" untilstart()finishesinitialSnapshot(). That function loads every enabled environment, both active and passive, throughloadEnvironmentShells, and it only returns once all of them have finished. Polling for the active computer (startPolling) doesn't start until after that.The shell requests run in parallel, each with the 6s timeout, so the splash stays up until the slowest one finishes. A peer that accepts the connection but never replies holds home for the whole timeout. A refused connection fails right away. That matches your table: about 2s when the port is closed, and about 8s when it accepts and stays silent.
There's one more wait inside that same group. When a passive environment has no cached config, its task fetches the catalogue after a successful shell, and on a cold start nothing is cached yet. Your hung run didn't hit this, because a failed shell returns before the catalogue request. I didn't measure a peer that's slow but succeeds.
Disabling the unresponsive environment avoids the wait, since
initialSnapshotonly loads enabled environments. I got that from the source and didn't retest it.Suggested fix
- Return the first snapshot as soon as the active environment's shell is ready, and start its polling then.
- Publish each passive environment when its own shell finishes. Until then, leave it unset rather than marking it disconnected early, and don't keep the loading screen up.
- If a passive shell fails or times out, keep any cached rows and mark that environment disconnected. A successful response should still apply when it arrives, no matter how long it took.
- Keep each passive environment's catalogue fetch inside that environment's own task.
- Leave
resumeAfterBackgroundas it is.
This fix belongs in
initialSnapshot. A fix for #14691 in the later refresh loop wouldn't change cold start, becauseinitialSnapshotstill waits for the full group.Related
#14691 (open issue) covers the passive refresh loop after launch. #10761 (closed PR) rewrote cold-start hydration along with passive workers and cache ownership. It was closed under the prior approval rule because that broader redesign had no triaged issue. This report covers just the cold-start wait, and accepting it doesn't mean adopting that rewrite.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 2, 2026
Area
apps/mobile: the native SwiftUI client in
apps/swift-iosont3code/rebuild-mobile-app-swift(#5178)Steps to reproduce
t3 serve, v0.0.44-nightly.20260929.2456, each with its own--base-dir): A on 127.0.0.1:47621 and B on 127.0.0.1:47622. Create one project and one thread on each throughPOST /api/orchestration/dispatch(project.create,thread.create).activeEnvironmentIDin the app'senvironments.json) and B is passive.simctl terminate, thensimctl launch) while recording the Simulator. Measure from the first launch frame to the first frame that shows home.Expected behavior
An unresponsive passive environment should not hide the active environment's threads at launch. Foreground resume already avoids this: "Other saved computers must not delay foreground recovery" (NativeFeatureClient.swift L280-L281).
Actual behavior
The app shows "Connecting to T3 Code" until B's shell request times out, then shows A's thread with B marked offline:
The app's network log for one hung launch shows A's two requests completing in 61ms and 22ms at 10:52:45.7. B's
GET /api/orchestration/shellfailed at 10:52:51.69 withNSURLErrorDomain -1001(timed out) after 6022ms. Home appeared only after that.From source:
initialSnapshot()awaitsloadEnvironmentShells(environments.filter(\.isEnabled))for every enabled environment before it returns (L228-L266, call at L243). Each shell read uses the 6s timeout (L163).FeatureRootModel.start()keeps the loading screen untilinitialSnapshot()returns (L123-L132).This is separate from #14691, which covers the passive refresh loop after launch and excludes cold start. A fix for #14691 inside
loadEnvironmentShellswould not change this on its own, becauseinitialSnapshot()still waits for the whole result.Impact
Minor bug or occasional failure
Version or commit
t3code/rebuild-mobile-app-swift@2eb6a53343(build 51). Its cold-start path (initialSnapshot,loadEnvironmentShells,startAggregateRefresh,FeatureRootModel.start, both timeouts) matches the current tipbb01adbf90, apart from 5 lines that pass failure details through. I couldn't build the tip here.T3Code.debug.dylibsha256e5bba723…). It shipped without entitlements, so the Keychain refused pairing (-34018). For this test I added only Xcode's Simulator entitlement section (application-identifier) to the small launcher executable and re-signed it ad hoc.t30.0.44-nightly.20260929.2456.Environment
iPhone 17 Pro Simulator, iOS 26.5, macOS host. Both servers and the non-replying listener ran on 127.0.0.1.
Logs or stack traces
Evidence files are on a fork pre-release. Each clip starts 0.5s before launch and ends 1.5s after home appears, and plays in real time.
Passive computer B accepts but never replies: "Connecting to T3 Code" for 7.8s, then home with B offline.
Control, B answers: home after 2.9s.
B refuses the connection: home after 2.4s.
Workaround
Disable the unresponsive environment before launching. This is from source (
initialSnapshotonly loads enabled environments); I did not test it.Question for triage
Is the complete first snapshot intended, waiting for every enabled environment before showing home? Or should cold start show the active environment as soon as its own shell loads, with passive environments filling in afterwards as resume already does? Closed #10761 rewrote this path together with the passive refresh and cache ownership. This report covers only the cold-start wait.
Investigated with Claude Opus 5.5 in Claude Code (T3 Code). Simulator taps were done by GPT-6 Astra (Codex) using the AXe CLI.