Skip to content

[Bug]: SwiftUI: one slow passive environment delays home updates from every other environment #14691

Description

@saphid

Area

apps/mobile — the native SwiftUI client in apps/swift-ios on t3code/rebuild-mobile-app-swift (#5178)

Steps to reproduce

  1. Run three local servers (t3 serve, v0.0.44-nightly.20260929.2456, each with its own --base-dir): A, B and C.
  2. In the SwiftUI app on an iOS Simulator, pair C, then B, then A, so that A is active and B and C are passive. B and C are reached through a local logging proxy, which can delay GET /api/orchestration/shell for C only.
  3. Leave the app on home. Rename a thread on B through B's own POST /api/orchestration/dispatch (thread.meta.update).
  4. Take a Simulator screenshot about every 0.9s and record the first frame where B's title changes. Compare that with the proxy's timing for B's and C's shell requests.
  5. Repeat with C answering immediately, C answering after 4.5s, and C not answering at all (the app gives up at its 6s shell timeout).

Expected behavior

A slow or unreachable passive environment should not hold back other environments. B's new title should appear as soon as B's own refresh returns.

Actual behavior

Passive environments are refreshed as one batch, and the batch publishes only after every environment has answered. B's rename reaches the app on time, but it stays off screen until C finishes or times out:

C behaviour Trials B title shown after B's response returned
Answers immediately 5 −0.7s to +0.0s (one outlier at +2.5s; screenshot resolution is ~0.9s)
Answers after 4.5s 3 +4.05s, +4.26s, +4.50s (C finished at +4.48s to +4.50s)
No answer (client timeout) 1 of 4 +5.74s (the app abandoned C at 6.0s)
No answer, C in its 20s back-off 3 of 4 −0.3s to +0.1s

B's refresh cadence also slows by the same amount:

  • C immediate: B was requested at 5.2s and 10.5s intervals.
  • C at 4.5s: 14.9s and 9.9s, so every cycle is ~4.5s later.
  • C not answering: 11.3s and 16.3s on cycles that include C.

A peer that answers just under the 6s timeout never enters back-off, so it delays every cycle indefinitely.

From source (not reproduced here): for a T3 Connect peer, the credential refresh runs before the 6s shell timeout applies (HTTP.swift, T3ConnectManagedAuthorization.swift). A slow relay could therefore hold the batch for longer than 6s. I have not measured this.

The batch await is in startAggregateRefresh and loadEnvironmentShells. Both functions are unchanged at the current branch tip bb01adbf90.

Impact

Minor bug or occasional failure

Version or commit

  • App: Debug build of t3code/rebuild-mobile-app-swift @ 285d2ab80c, plus unrelated review-view changes; the refresh code is identical to the tip. It was re-signed ad hoc with only its app-group entitlement so pairing could persist on the Simulator.
  • Server: t3 0.0.44-nightly.20260929.2456.

Environment

iPhone 17 Pro Simulator, iOS 26.5, macOS host. All servers and proxies ran on 127.0.0.1.

Logs or stack traces

Evidence files are on a fork pre-release.

Trial 5 (C answers after 4.5s), before and after B's rename. The first changed screenshot came 13.82s after the rename. B's own response had returned at 9.31s and C's at about 13.8s. Screenshots were about 0.9s apart, so each time can be up to ~0.9s late.

B thread v4 before and v5 after, trial 5

  • Per-trial timing and provenance receipt (JSON) — every trial and phase, the uncertainty notes, the build provenance (not a build of bb01adbf90), and asset hashes.
  • Supplemental recording (MP4) — supplemental only. It is a real Simulator recording of a separate trial with C at 4.5s, and it shows the title change. The Simulator records frames only when the screen changes, and the recording's start offset relative to the rename is unknown, so it does not demonstrate timing.

Workaround

Disable the slow environment.

Question for triage

Is this the intended behaviour? If not, would a bounded per-environment deadline inside the existing batch be the expected fix? That would let the batch publish without waiting past the deadline, while the slow environment keeps its cached rows and uses the existing back-off. Or do you want independent per-environment refresh? Closed #10732 took the latter approach and was declined for scope. Cold-start hydration (#10761) is a separate question and is not covered here.

Investigated with Claude Opus 5.5 in Claude Code (T3 Code).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions