Skip to content

[Bug]: SwiftUI: an unresponsive passive environment delays cold start of the active one #14699

Description

@saphid

Area

apps/mobile: the native SwiftUI client in apps/swift-ios on t3code/rebuild-mobile-app-swift (#5178)

Steps to reproduce

  1. Run two local servers (t3 serve, v0.0.44-nightly.20260929.2456, each with its own --base-dir): A on 127.0.0.1:47621 and B on 127.0.0.1:47622. Create one project and one thread on each through POST /api/orchestration/dispatch (project.create, thread.create).
  2. In the SwiftUI app on an iOS Simulator, pair B, then A. A is the active environment (activeEnvironmentID in the app's environments.json) and B is passive.
  3. Make B unresponsive without removing it. Stop B and listen on its port with a socket that accepts TCP connections but never replies. This stands in for a paired computer that is asleep or stuck behind a stalled network path.
  4. Cold-launch the app (simctl terminate, then simctl launch) while recording the Simulator. Measure from the first launch frame to the first frame that shows home.
  5. Repeat with both servers answering (control) and with B's port closed (connection refused).

Expected behavior

An unresponsive passive environment should not hide the active environment's threads at launch. Foreground resume already avoids this: "Other saved computers must not delay foreground recovery" (NativeFeatureClient.swift L280-L281).

Actual behavior

The app shows "Connecting to T3 Code" until B's shell request times out, then shows A's thread with B marked offline:

B behaviour Launch → home (3 cold launches each)
Answers (control) 2.93s, 2.27s, 1.73s
Accepts but never replies 8.84s, 7.84s, 7.83s
Port closed (refused) 2.44s, 2.42s, 2.38s

The app's network log for one hung launch shows A's two requests completing in 61ms and 22ms at 10:52:45.7. B's GET /api/orchestration/shell failed at 10:52:51.69 with NSURLErrorDomain -1001 (timed out) after 6022ms. Home appeared only after that.

From source: initialSnapshot() awaits loadEnvironmentShells(environments.filter(\.isEnabled)) for every enabled environment before it returns (L228-L266, call at L243). Each shell read uses the 6s timeout (L163). FeatureRootModel.start() keeps the loading screen until initialSnapshot() returns (L123-L132).

This is separate from #14691, which covers the passive refresh loop after launch and excludes cold start. A fix for #14691 inside loadEnvironmentShells would not change this on its own, because initialSnapshot() still waits for the whole result.

Impact

Minor bug or occasional failure

Version or commit

  • App: Debug Simulator build of t3code/rebuild-mobile-app-swift @ 2eb6a53343 (build 51). Its cold-start path (initialSnapshot, loadEnvironmentShells, startAggregateRefresh, FeatureRootModel.start, both timeouts) matches the current tip bb01adbf90, apart from 5 lines that pass failure details through. I couldn't build the tip here.
  • That build's code is byte-identical to its build receipt (T3Code.debug.dylib sha256 e5bba723…). It shipped without entitlements, so the Keychain refused pairing (-34018). For this test I added only Xcode's Simulator entitlement section (application-identifier) to the small launcher executable and re-signed it ad hoc.
  • Server: t3 0.0.44-nightly.20260929.2456.

Environment

iPhone 17 Pro Simulator, iOS 26.5, macOS host. Both servers and the non-replying listener ran on 127.0.0.1.

Logs or stack traces

Evidence files are on a fork pre-release. Each clip starts 0.5s before launch and ends 1.5s after home appears, and plays in real time.

Passive computer B accepts but never replies: "Connecting to T3 Code" for 7.8s, then home with B offline.

Hung passive computer: Connecting to T3 Code for 7.8s before the active computer's thread appears

Control, B answers: home after 2.9s.

Control: both computers answer and home appears 2.9s after launch

B refuses the connection: home after 2.4s.

Refused connection: home appears 2.4s after launch

Workaround

Disable the unresponsive environment before launching. This is from source (initialSnapshot only loads enabled environments); I did not test it.

Question for triage

Is the complete first snapshot intended, waiting for every enabled environment before showing home? Or should cold start show the active environment as soon as its own shell loads, with passive environments filling in afterwards as resume already does? Closed #10761 rewrote this path together with the passive refresh and cache ownership. This report covers only the cold-start wait.

Investigated with Claude Opus 5.5 in Claude Code (T3 Code). Simulator taps were done by GPT-6 Astra (Codex) using the AXe CLI.

Activity

  1. juliusmarminge commented on Oct 2, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the cold-launch timings and the clear comparison between refused and silent connections, @saphid! This is a real bug. I confirmed it on t3code/rebuild-mobile-app-swift at bb01adbf90, the branch tip you named (open PR #5178). It only affects the experimental SwiftUI client. The React Native app we ship doesn't use this startup path.

    To answer your question: cold start should show the active environment as soon as its own shell is ready, and passive environments should fill in afterwards. Waiting for every enabled environment before leaving "Connecting to T3 Code" isn't the intended behavior. That matches foreground resume, which already says other saved computers must not delay recovery.

    What happens today

    FeatureRootModel.isLoading starts out true, and FeatureRootView shows "Connecting to T3 Code" until start() finishes initialSnapshot(). That function loads every enabled environment, both active and passive, through loadEnvironmentShells, and it only returns once all of them have finished. Polling for the active computer (startPolling) doesn't start until after that.

    The shell requests run in parallel, each with the 6s timeout, so the splash stays up until the slowest one finishes. A peer that accepts the connection but never replies holds home for the whole timeout. A refused connection fails right away. That matches your table: about 2s when the port is closed, and about 8s when it accepts and stays silent.

    There's one more wait inside that same group. When a passive environment has no cached config, its task fetches the catalogue after a successful shell, and on a cold start nothing is cached yet. Your hung run didn't hit this, because a failed shell returns before the catalogue request. I didn't measure a peer that's slow but succeeds.

    Disabling the unresponsive environment avoids the wait, since initialSnapshot only loads enabled environments. I got that from the source and didn't retest it.

    Suggested fix

    • Return the first snapshot as soon as the active environment's shell is ready, and start its polling then.
    • Publish each passive environment when its own shell finishes. Until then, leave it unset rather than marking it disconnected early, and don't keep the loading screen up.
    • If a passive shell fails or times out, keep any cached rows and mark that environment disconnected. A successful response should still apply when it arrives, no matter how long it took.
    • Keep each passive environment's catalogue fetch inside that environment's own task.
    • Leave resumeAfterBackground as it is.

    This fix belongs in initialSnapshot. A fix for #14691 in the later refresh loop wouldn't change cold start, because initialSnapshot still waits for the full group.

    Related

    #14691 (open issue) covers the passive refresh loop after launch. #10761 (closed PR) rewrote cold-start hydration along with passive workers and cache ownership. It was closed under the prior approval rule because that broader redesign had no triaged issue. This report covers just the cold-start wait, and accepting it doesn't mean adopting that rewrite.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions