Skip to content

[Bug]: Orchestrator V2 event writes stall the server on a high-latency disk (156 ms fsync), environment flips Offline #15707

Description

@deepontology

Before submitting

Area

apps/server

Steps to reproduce

  1. Run 0.0.46-nightly.20261004.2644 with ~/.t3/userdata on a disk with high fsync latency. On this machine that is the root disk: Toshiba HDWD110 (7200 RPM), crypto_LUKS, btrfs with compress=zstd:3. Measured fsync: 156 ms average over 200 4K writes. The same test on tmpfs: ~0 ms.
  2. Start several concurrent runs. I had four opencode runs and one claudeAgent session.
  3. After about an hour the desktop and Android clients cycle between Connected and "connect to the environment".

Expected behavior

The server keeps answering keepalives while agents stream, or degrades without dropping clients.

Actual behavior

Orchestrator V2 (the OrchestrationV2 migration; tables prefixed orchestration_v2_) persists each provider stream delta as an event. The server process sat in uninterruptible I/O wait for long stretches. Sampled wchan values: btrfs_btree_wait_writeback_range, write_all_supers, folio_wait_bit_common, wait_log_commit. SQLite transactions reached 1.98 s. Trace segments were dominated by orchestrationV2.EventSink.write spans of 40 to 500 ms each. Clients reported:

The server log shows repeated SessionStore.markConnected and markDisconnected cycles. The Android client stayed logged in and showed the same offline state.

Event volume: 10,835 events were appended between 15:50 and 16:15 UTC, about 5,500 in the worst ten minutes, so roughly nine writes per second. In the last 5,000 rows there were 416 distinct node IDs, and one node ID appeared 405 times within 2,000 rows. This matches the mechanism in #14701: one synchronous SQLite connection on the main thread, so a slow commit blocks every other request, and each reconnect repeats the work.

Earlier builds with several concurrent runs did not show this on the same machine.

Impact

Major degradation or frequent failure. T3 Code is unusable while agents run on this machine.

Version or commit

0.0.46-nightly.20261004.2644 (Arch package t3code-nightly-bin 0.0.46_nightly.20261004.2644-1)

Environment

Arch Linux, Wayland. Root and /home on crypto_LUKS + btrfs over a Toshiba HDWD110 HDD, compress=zstd:3. Desktop and Android clients on the same account.

What I tried

  • Moved ~/.t3/userdata to a btrfs directory with chattr +C (no COW, no compression). No change. The cost is commit latency on the disk, not COW.
  • Aborted the runs through the opencode API. Each parent recorded MessageAbortedError, and the server kept inserting node.updated rows for about a minute.
  • Terminated the provider processes and restarted the app. Stalls returned with the next concurrent runs.

Workaround

Reverting to the previous nightly, 0.0.45_nightly.20260930.2468, which predates the Orchestrator V2 event store. state.sqlite is intact, so this restores normal operation on this hardware. I will update after running it for a while.

Open questions

Logs or stack traces

Happy to attach server trace segments, the wchan samples, and the fsync benchmark script if useful.

Related

#14701, #15609, #15442, #14871

Submitted with an opencode agent 🤖

Activity

  1. juliusmarminge commented on Oct 4, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the fsync numbers, trace spans, and event counts, @deepontology. That made this very concrete.

    What I found

    I checked v0.0.46-nightly.20261004.2644 and current main (4ee6bfd50e). The write path is the same in both. The only SQLite change since your nightly is #15488, which switches transactions to BEGIN IMMEDIATE but still commits on the same connection.

    • apps/server/src/persistence/Layers/Sqlite.ts opens a single node:sqlite DatabaseSync in WAL mode. It never sets PRAGMA synchronous, so SQLite runs at FULL and fsyncs the WAL on every commit.
    • packages/shared/src/nodeSqliteClient.ts puts that connection behind a one-permit semaphore and runs statements and commits on the main thread. A slow fsync therefore blocks the event loop and every other statement.
    • Each provider event is committed on its own. ProviderEventIngestor.ingestNormalized calls EventSink.write, which appends the event, applies projections, and commits in one transaction. At about 156 ms per fsync, the disk can do roughly six commits a second. Your stretches at about nine events a second line up with the I/O wait you saw.
    • Assistant text is already coalesced in the default paragraph streaming mode. Two paths still write on every update:
      • OpenCode emits a node.updated for each reasoning delta, and those are not coalesced like Claude reasoning and Codex text.
      • Tool updates (node.updated and turn_item.updated) have no streaming filter.
    • That fits a single node ID appearing hundreds of times in a short window. After an abort, events that are already queued still commit one at a time until the queue empties.
    • The client errors follow from the stall. The environment descriptor fetch has a 10 s timeout, and the websocket closes after missed pongs. Each reconnect cycles SessionStore.markDisconnected / markConnected and reloads the snapshot.

    Likely fix area

    Some options:

    • Use PRAGMA synchronous=NORMAL in WAL mode, which syncs at checkpoint instead of on every commit.
    • Coalesce OpenCode reasoning and tool updates, and batch queued provider events into one transaction.
    • Move the writer off the event loop so a slow commit can't starve keepalives.

    #14701 is the read-path side of the same single connection. Open PR #14703 moves that read to a worker but leaves writes on the main thread, so it wouldn't cover this case. A trace slice from a stall would still help show the mix of events.

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 4, 2026
  3. deepontology commented on Oct 4, 2026

    @deepontology
    Author

    Trace data from a second stall, as requested.

    Setup: same machine and build as the report. Two opencode runs started at 16:32 UTC, and the environment flapped again while they streamed. The trace files cover 16:38 to 17:13 UTC.

    Event mix, 16:32 to 17:13 UTC, from orchestration_events:

    event type count
    node.updated 18,171
    turn-item.updated 2,847
    provider-thread.updated 490
    message.updated 163
    provider-turn.updated 52

    node.updated by kind: reasoning 16,549 (91.1%), tool_call 1,448 (8.0%), assistant_message 158 (0.9%), root_turn 12.

    By provider: opencode 18,116 (99.7%), claudeAgent 55.

    Most-updated native items, all OpenCode reasoning nodes:

    • prt_107c694da001... 1,049 updates
    • prt_107cdb427001... 715 updates
    • prt_107ca6c7a001... 638 updates
    • prt_107cf4ab4001... 579 updates

    Worst minute, 17:10:00 to 17:11:00 UTC, from server.trace.ndjson:

    • orchestrationV2.EventSink.write: 487 spans, 106,295 ms total
    • sql.transaction: 514 spans, 108,681 ms total
    • 106.3 s of commit time in a 60 s minute is 177% saturation on the single-threaded writer
    • durations: p50 100 ms, p95 1,125 ms, max 3,314 ms
    • busy time per second stayed mostly above 1,000 ms, so the queue never drained

    The pattern holds across every minute of the incident. All 37 minutes in the trace show 104 to 112 s of transaction time per minute.

    The first incident showed the same mix: 15:50 to 16:15 UTC had 7,619 node.updated, of which 6,022 (79%) were reasoning and 1,475 (19.4%) were tool_call, with opencode responsible for 7,554 (99.1%).

    Artifacts in this gist: https://gist.github.com/deepontology/ac5bdb621948c59bebde99efae4768da

    • trace-slice.jsonl: every span in the worst minute, fields are ts, span, ms, status, trace. 4,501 rows. Attributes, IDs, and payloads stripped.
    • per-minute.csv: sink and transaction counts and busy time for 16:38 to 17:13 UTC.
    • per-second-worst-minute.csv: per-second commit count and busy time for 17:10 UTC.

    This matches your triage. OpenCode reasoning deltas dominate, that path has no coalescing, and the commits serialize on the one connection.

    Submitted with an opencode agent 🤖

  4. AnalogCyan commented on Oct 5, 2026

    @AnalogCyan

    Another data point, on fast storage, focused on how much this writes rather than fsync latency.

    Setup: t3 serve 0.0.46-nightly.20261005.2689 as a systemd user service, Debian 13, Crucial P3 NVMe, btrfs (compress=zstd:3), 16 GB RAM. Mostly Claude agents, often 5 to 10 at once, plus Codex delegated tasks.

    statev2.sqlite is 1.3 GB with 0 free pages. Growth of orchestration_events by day (MB of payload_json):

    Day Events MB
    09-20 to 10-02, typical 2k to 18k 1 to 15
    10-03 (V2 in use from here) 36,883 84
    10-04 40,742 151
    10-05 39,984 114

    By type, all time:

    event_type count MB
    turn-item.updated 46,778 250
    thread.activity-appended 64,191 113
    node.updated 29,967 31
    provider-turn.updated 13,084 15

    turn-item.updated stores the whole item on each update, so long tool outputs are written again on every update. Single orchestration_v2_projection_turn_items rows reach 680 KB, and that table is another 200 MB. One long-running thread with its subagents holds 264 MB of events and items. Nothing prunes events, including those of archived threads.

    Read side (#14701 territory), from the trace: a message.dispatch transaction took 6.1 s, of which 3.4 s was SELECT payload_json FROM orchestration_v2_projection_turn_items WHERE thread_id = ? for a thread of 1,769 items (3.4 MB). The host was under swap pressure, so pages were cold, but the query loads every item of the thread on the main thread. The server logs "event loop stalled" of 2 to 8 s several times a day, once 41 s, and remote clients drop and resync on each.

    Coalescing tool turn_item.updated writes as you suggested would cut most of this volume. A retention rule for events of archived or old threads would bound the rest; I can open an Ideas discussion for that if it helps.

  5. apra31 commented on Oct 6, 2026

    @apra31

    I’m seeing similar Offline/Reconnecting symptoms on Windows desktop with a WSL2 backend, using a newer V2 nightly. I captured the server’s main thread waiting in fsync() for an ext4 journal commit while T3 was showing Offline/Reconnecting.

    Environment

    • Desktop and backend: 0.0.46-nightly.20261005.2702

    • WSL distribution: Ubuntu 24.04

    • “WSL only” enabled

    • Local backend: http://127.0.0.1:3773

    • ~/.t3 resolves inside the WSL home directory

    • Filesystem: /dev/sdf, ext4, mounted at /

    • Mount options: rw,relatime,discard,errors=remount-ro,data=ordered

    Filesystem capacity:

    Filesystem  Size   Used  Avail  Use%  Mounted on
    /dev/sdf    1007G   386G   570G   41%  /

    Physical drive details and Windows host free space have not been collected. I have not independently verified the live database’s exact path.

    Workload and symptoms

    I started a coordinator with four ordinary worker threads through T3’s thread-management tools. Multiple Flutter compilation/test processes were running as part of their work.

    The environment showed Offline/Reconnecting while Claude and build/test processes remained alive. This is the first occurrence I have observed in my setup.

    Process and I/O evidence

    At approximately 18:34 UTC+7 on October 6, the T3 server process showed:

    PID   PPID  STAT  WCHAN                  COMMAND
    1674  1673  Dsl+  folio_wait_bit_common  node-MainThread

    Six visible Claude processes had PID 1674 as their parent.

    The four interval samples from vmstat 1 5, excluding the first since-boot report, showed:

    Sample | Blocked processes (b) | I/O wait (wa) | CPU idle (id) -- | -- | -- | -- 1 | 10 | 38% | 53% 2 | 8 | 24% | 61% 3 | 7 | 27% | 67% 4 | 8 | 27% | 59%

    Swap usage was approximately 1.4 GiB. Three Dart frontend processes held approximately 4.65 GiB of resident memory combined.

    The server still owned port 3773. Its listening socket’s Recv-Q was 7 at 18:34 and 1 at 18:44.

    Kernel stack captured while T3 was still Offline/Reconnecting

    At approximately 18:44–18:45 UTC+7:

    sudo cat /proc/1674/stack
    

    [<0>] jbd2_log_wait_commit+0xb4/0x130
    [<0>] jbd2_complete_transaction+0x73/0xb0
    [<0>] ext4_fc_commit+0x64b/0x9f0
    [<0>] ext4_sync_file+0x1e4/0x2b0
    [<0>] vfs_fsync_range+0x49/0x90
    [<0>] do_fsync+0x40/0x90
    [<0>] __x64_sys_fsync+0x17/0x20
    [<0>] x64_sys_call+0x1729/0x20f0
    [<0>] do_syscall_64+0x73/0x990
    [<0>] entry_SYSCALL_64_after_hwframe+0x76/0x7e

    This captures a synchronous filesystem commit wait in the server’s main thread. It does not identify the file being synchronized or measure the duration of that operation, so I cannot yet confirm that SQLite is the caller or that this has the same root cause.

    Would specific T3 traces or further storage diagnostics help distinguish the V2 event-write path from WSL storage contention caused by concurrent builds/tests?

  6. PiquelChips commented on Oct 7, 2026

    @PiquelChips

    Another data point: an NVMe SSD on ext4, where the trigger is other processes writing heavily rather than a slow disk. On SSH-managed environments the stall also ends up killing running agents.

    Setup

    • 0.0.46-nightly.20261007.2774, started by the desktop app as an SSH-managed environment (~/.t3/ssh-launch/…/run-t3.sh)
    • NixOS, kernel 6.18.55, 58 GB RAM, no swap
    • ~/.t3 on ext4 (default data=ordered), Crucial P3 4 TB NVMe (CT4000P3SSD8, firmware P9CR30A)
    • statev2.sqlite is 2.3 GB
    • Workload: several Claude agents, each building a Rust workspace in its own worktree (cargo build / cargo nextest, debuginfo=2)

    Numbers

    • fsync of a 4 KB file in ~/.t3/userdata: 26–45 ms normally, spiking to 5.5 s while the builds were writing about 60 MB/s.
    • At the worst point there were 8.7 GB of dirty pages and /proc/pressure/io showed full avg10=81. Memory pressure was 0, so tools like btop showed nothing wrong.
    • server.eventLoop.stall reached 15–16 s several times. CPU use during the stalls was low (utilization 0.2–0.5, a few tens of ms of CPU), so the main thread was waiting, not computing.
    • Trivial statements took seconds: UPDATE auth_sessions SET last_connected_at = ? took 4.2 s and SELECT * FROM scheduled_tasks WHERE enabled = 1 … took 7.5 s. One sql.transaction took 23.6 s.
    • The only PRAGMAs in the traces are journal_mode = WAL, busy_timeout = 5000, foreign_keys = ON and journal_size_limit. That matches triage: synchronous is left at FULL.

    With ext4 data=ordered, a journal commit can wait for other files' dirty data to be written out first. So a small WAL fsync gets queued behind the build output, even on an SSD. This probably answers the WSL comment's question: concurrent builds and tests are enough to trigger it without a slow disk.

    What it does to SSH-managed environments

    The stall gets worse than "Offline" here:

    1. The event loop blocks for 10–16 s.
    2. The desktop's SSH session drops, the client reconnects, and the launcher starts a new server. That happened 6 times in about 20 minutes.
    3. Startup recovery then ends whatever was running, e.g. V2 orchestration recovery completed { terminalizedRuns: 3, stoppedSessions: 4 } at 18:22:51.

    Each restart kills in-flight agent turns and their child processes. Inside the agents the builds show up as exit 137, and the agents reported this to me as "out of memory", even though there was no kernel or oomd kill and about 50 GB of RAM was free.

    A second t3 serve on the same ~/.t3 (the #5749 / #6097 situation) makes this worse. Each slow commit holds the database lock for seconds, and the other server fails with LockTimeoutError: database is locked once busy_timeout runs out.

    Mitigations on my side

    Fewer concurrent builds, lower vm.dirty_bytes / vm.dirty_background_bytes, and debug = "line-tables-only" for dev builds. These shrink the spikes but don't remove them. PRAGMA synchronous=NORMAL in WAL mode, or moving the writer off the main thread, as suggested in triage, would fix it at the source.

    Trace segments covering the stalls and restarts are available if useful.


    Investigated and written by Claude Opus 5.5 running in Claude Code (via T3 Code).

  7. svenpeeters commented on Oct 7, 2026

    @svenpeeters

    Another data point: the same disconnect cycle on fast storage. My fsync latency is about 30 times lower than in the original report, so slow disks aren't the only trigger.

    Setup

    • 0.0.46-nightly.20261005.2702 (Arch, t3code-nightly-bin), desktop app, local environment only
    • Omarchy (Arch), kernel 7.2.5, 32 cores, 125 GB RAM
    • ~/.t3 on btrfs (compress=zstd:3) on LUKS, Samsung 990 PRO 2 TB NVMe
    • fsync of a 4 KB write in ~/.t3/userdata: 5.0 ms average (p50 4.9 ms, max 5.6 ms, 50 samples)
    • Workload: Claude agents only (claudeAgent)

    Symptoms

    The add-project flow showed "Environment unavailable – omarchy is not connected", and the connection kept dropping and coming back. In server.trace.ndjson, top-level EnvironmentRpc.subscribe spans were cut short again and again between 18:03 and 18:13 UTC, each lasting 5 to 133 seconds (8.4 s, 16 s, 48 s, 93 s, 133 s, 60 s, …). The busiest minute of event writes, 18:11 UTC with 719 events, falls inside that window. The rotated server trace files filled up quickly: 10 MB in about 4 minutes, mostly sql.execute, db.transaction.commit, orchestrationV2.EventSink.write and ThreadSettlementServiceV2.settleThread.

    The desktop process had been running for 7h40m with a 9.6 GB memory peak (reported by systemd for the app's scope). A full relaunch fixed it for now.

    Database

    • statev2.sqlite: 564 MB, 0 free pages. The legacy state.sqlite (253 MB) is still on disk.
    • Largest tables: orchestration_events 265 MB, orchestration_v2_projection_turn_items 101 MB, projection_thread_activities 91 MB
    • orchestration_v2_events is empty. Every event is written to orchestration_events.

    Events by type, all time:

    event_type count MB
    turn-item.updated 12,046 99.5
    thread.activity-appended 20,567 89.5
    node.updated 6,720 6.0
    message.updated 4,166 3.3
    provider-turn.updated 3,960 3.6

    turn-item.updated averages about 8.3 KB per event, which looks like the full item rewritten on every update rather than only what changed.

    Daily growth:

    Day Events MB
    10-05 639 0.6
    10-06 17,528 86.2
    10-07 (partial) 7,991 29.2

    With fast storage, disk speed doesn't explain the stalls. The large database, plus the read-path blocking from #14701, seems to be enough to block the main thread even at 5 ms per commit.

  8. Sypher760-gif commented on Oct 9, 2026

    @Sypher760-gif
    Contributor

    I have a small patch ready for the first option in the triage comment: PRAGMA synchronous = NORMAL in layerSetup (apps/server/src/persistence/Sqlite.ts), next to the other pragmas. It is one line plus a test that the pragma is applied.

    Before opening a PR, can a maintainer confirm this direction? It changes a default: in WAL mode, NORMAL stops fsyncing the WAL on every commit. After an OS crash or power loss the last few committed transactions can be lost. An application crash is safe, and WAL with NORMAL cannot corrupt the database. Checkpoints still fsync.

    A quick measurement on an NVMe disk (Windows, node:sqlite, WAL, 1000 single-row transactions): 0.431 ms per commit at FULL, 0.050 ms at NORMAL. That is a fast disk, so it does not reproduce the 156 ms case from the report. I can't confirm that this alone stops the Offline flapping, and it does not touch event volume or moving the writer off the main thread.

    If this direction is fine, I'll open the PR with disk-backed timings. If you would rather go another way, I'll leave it.

  9. iamparmjeet commented on Oct 9, 2026

    @iamparmjeet

    Another data point on fast storage (NVMe + btrfs + LUKS), with disk-backed FULL vs NORMAL timings from the same disk.

    Setup

    • 0.0.46_nightly.20261006.2735 (Arch, t3code-nightly-bin), desktop app, local environment only
    • Omarchy, kernel 7.2.8, 16 GB RAM
    • ~/.t3 on btrfs (compress=zstd:3) on LUKS, XPG GAMMIX S70 BLADE 1 TB NVMe (SMART clean, 4% used, 0 media errors)
    • Workload: Claude, Codex and OpenCode, with delegated tasks across providers

    Stall evidence

    A watcher polled /.well-known/t3/environment every 30 s and captured system state when it failed. It recorded 11 stalls, each a 5 s timeout, while the UI showed reconnecting:

    • The server main thread was in D state with wchan folio_wait_bit or wait_log_commit.
    • /proc/pressure/io showed some avg10 between 17 and 42, while CPU and memory pressure were about 0.
    • In 8 of the 11 snapshots, the T3 server was the only userspace process in D state.

    During one stall (16:50 to 16:53 local time), queries that are normally fast stretched out:

    span median max
    SELECT * FROM scheduled_tasks … 0 ms 23.4 s
    getLimitRecoveryCandidates 3 ms 23.4 s
    orchestrationV2.EventSink.write — 1 to 6 s
    client ConnectionDriver.connect — 10 s

    Write volume

    • With agents running, the server process wrote 64 MB to disk in 30 s, across 26,598 write syscalls (/proc/<pid>/io).
    • statev2.sqlite is 2.5 GB. orchestration_events holds 1.46 GB in 467k rows.
    • payload_json per day: 10-05 39 MB, 10-06 42 MB, 10-07 161 MB, 10-08 230 MB.
    • 10-08 to 10-09 by type: turn-item.updated 39,535 rows (182 MB), node.updated 88,396 rows (141 MB).

    NOCOW does not fix it

    This confirms the original report. Rewriting statev2.sqlite as a chattr +C file cut it from 157,192 extents to 92. Stalls returned within 40 minutes, with the same wchan values.

    FULL vs NORMAL on this disk

    Test conditions: node:sqlite on Node v26.11.1, SQLite 3.53.4, WAL, file on the same btrfs/LUKS volume as ~/.t3/userdata. Each run is 1000 single-row transactions (BEGIN IMMEDIATE + insert of a ~4.6 KB payload + COMMIT), roughly one turn-item.updated each. The disk was idle.

    p50 p99 max total
    raw 4 KB write + fsync 1.56 ms 2.12 ms (p95) 2.26 ms —
    synchronous=FULL 1.71 ms 3.03 ms 6.66 ms 1.75 s
    synchronous=NORMAL 0.03 ms 0.13 ms 8.80 ms 0.08 s

    On an idle disk, FULL already costs about one fsync per event. Under load the stall snapshots show those waits growing to seconds, and every one happens on the main thread. NORMAL removes the per-commit wait, leaving only checkpoint syncs. That is the max outlier in the NORMAL row.

    @Sypher760-gif, happy to run your patch on this machine and report whether the stalls stop under the same workload.

  10. MaximilianMauroner commented on Oct 9, 2026

    @MaximilianMauroner

    Note

    gpt-6.1-sol · Codex running in T3Code, responding on behalf of Max.

    Another occurrence on an NVMe-backed Linux VM, with 104–122 second main-thread stalls during one-event V2 writes, plus further stalls after the competing emulator workload stopped.

    This supports the responsiveness failure tracked here. It does not establish that T3 caused the underlying storage slowdown or that the SSD is defective.

    Observed behavior and impact

    On 2026-10-09 the macOS client showed:

    T3 Connect · Reconnecting: Could not connect relay environment.

    The server process remained alive, but even local HTTP requests timed out. Remote access recovered without a restart when disk pressure fell, then additional long stalls appeared. This prevented reliable interaction with running work in the remote environment.

    Expected: storage contention can delay persistence, but a blocked database operation should not freeze the main HTTP/WebSocket event loop and cause the entire environment to disappear behind a generic relay error.

    Environment

    • Running server: 0.0.46-nightly.20261009.2861, Linux x64, systemd user service; kernel 6.8.0-146-generic. The matching release tag points to 3b6af0bd1600f034b0466e1b8ff017fbabeba910.
    • Remote connection: T3 Connect with managed cloudflared 2026.10.0, accessed from the macOS desktop app. The desktop version was not collected.
    • Linux VM: 6 vCPUs, 6 GB allocated RAM (5,919 MiB visible), ext4 on LVM; root filesystem about 94% used with about 31 GB available during diagnosis. About 2 GB of swap was occupied during the severe incident. Swap occupancy alone does not establish active swapping as the cause.
    • Live V2 database: ~/.t3/userdata/statev2.sqlite, approximately 12 GiB by rounded ls -lh at 13:03 UTC; old state.sqlite approximately 7.4 GiB. The V2 WAL was approximately 13 MiB at that snapshot. These are file sizes, not measured event payload totals or growth rates.
    • Storage route: guest ext4/LVM → virtio raw virtual disk (cache=writeback, discard=unmap) → Unraid /mnt/user FUSE (fuse.shfs) → ZFS cache pool → Crucial CT500P2SSD8 NVMe SSD. The image physically resides on the cache SSD pool. No explicit QEMU IOThread was configured.
    • Hypervisor: about 16 GB RAM total; sampled available memory about 3 GB, ZFS ARC about 2 GB. These host measurements were after the worst stall.

    Correlated trace evidence

    All times below are UTC on 2026-10-09. The first column identifies the end of the correlated activity, not its start. Durations come from T3's trace; stall duration is the event-loop span's delayMaxMs attribute, not the tiny duration of the monitoring span itself.

    End time sql.transaction orchestrationV2.EventSink.write Event count on write Event-loop delay
    12:20:05 104,464.394 ms 104,467.450 ms 1 104,495 ms
    12:23:21 122,614.470 ms 122,616.425 ms 1 122,480 ms
    12:57:09–10 69,038.483 ms and a separate 44,107.001 ms transaction 44,108.285 ms 1 69,800 ms

    The later row contains overlapping spans, not one transaction attributed wholly to one event. A sql.execute SELECT span ending at 12:57:09 lasted 69,060.493 ms. Its elapsed time can include waiting for the shared connection, so this does not establish a 69-second query execution cost.

    The monitor reported relatively little CPU use over the long sampled intervals:

    Stall User CPU System CPU RSS Major page faults
    104,495 ms 2,501 ms 2,544 ms 506 MiB 158
    122,480 ms 785 ms 727 ms 535 MiB 28
    69,800 ms 1,867 ms 1,193 ms 548 MiB 56

    These counters cover the monitor's sample interval, not an isolated native SQLite call. High event-loop utilization here must not be read as high CPU utilization.

    A redacted scan of retained trace files, collected at 13:06 UTC and covering span endings from 12:18:04 to 13:05:15, contained:

    • 23 server.eventLoop.stall records.
    • 7,100 sql.transaction spans.
    • 4,591 orchestrationV2.EventSink.write spans.
    • 43,258 sql.execute spans.
    • 9,446 ThreadSettlementServiceV2.settleThread spans; several reached about 123 seconds during the severe stall.

    These are trace span counts, not unique database events, physical write counts, or additive time on disk. The retained window is incomplete before 12:18 due to log rotation. Long-lived subscription spans were excluded from interpreting request latency.

    Process, kernel, HTTP and relay evidence

    During the severe incident:

    • The server's main thread was sampled in D state, then with wait channel jbd2_log_wait_commit, an ext4 journal-commit wait. No complete native stack or syscall-to-file attribution was captured.
    • /proc/pressure/io: some avg10=96.47%, full avg10=87.10%.
    • vmstat interval samples: 47–61% CPU I/O wait, six blocked tasks; load average 33.14 on six vCPUs.
    • HTTP on 127.0.0.1:3773 timed out after four seconds. The public environment endpoint timed out after five seconds from the VM and fifteen seconds from the Mac.
    • Cloudflared reported canceled /api/t3-connect/health and /api/t3-connect/mint-credential requests and QUIC inactivity timeouts. Local HTTP failure means this was not solely a relay/network outage.
    • No relevant kernel warning was found in the checked 20-minute window.

    Around 12:24:24, without restarting or changing configuration, local HTTP returned 200 in 83 ms and the public endpoint returned 200 in 163 ms. I/O PSI fell to some avg10=1.10%, full avg10=0.87%; the connector then reported connected/fresh. This was temporary recovery, not a demonstrated fix.

    Competing workload and storage health

    An Android emulator started at 12:11:47, used about 2.3 GB RSS, and /proc/<pid>/io showed about 17.1 GB written after roughly twelve minutes. Its QA workload deliberately tested Android guest storage exhaustion using approximately 5.6 GB filler writes and flushes.

    The fault was confined to the emulator's /data virtual filesystem. The probe retained a 24 GiB reserve on the Linux VM filesystem. The final probe recorded 118,784 bytes free inside Android before the test, then removed the filler and restored approximately 5.6 GB free. This was not evidence that T3's filesystem ran out of space. The workload is strongly correlated with the early incident but was not isolated in a controlled A/B test.

    The emulator was absent by the 12:40 check and again at the follow-up process check. The 69.8-second stall at 12:57 therefore shows that stopping it did not eliminate all pauses. Remaining workload, delayed writeback, database read cost, memory pressure, and storage queueing have not been separated.

    After recovery, SSD SMART showed PASSED, critical warning 0, zero media/data-integrity errors, 46°C and 5% used. ZFS was ONLINE with no known data errors; the pool had about 154 GB free. Historical ZFS averages showed approximately 237 ms total write wait versus 3 ms device write wait and 604 ms synchronous queue wait, but these were not measurements from the exact outage. Idle follow-up samples were about 2 ms total write wait and very low NVMe utilization. This suggests investigating queueing through the storage stack; it does not prove FUSE caused the stall or rule out SSD latency under load.

    Relevant code in this exact release

    I checked the matching release source, rather than relying only on current main:

    1. nodeSqliteClient.ts: connection construction opens NodeSqlite.DatabaseSync. Statement execution directly invokes statement.all() / statement.run() inside the Effect runtime. An Effect wrapper does not move these native synchronous calls to another thread.
    2. The same adapter uses a one-permit semaphore around one connection and keeps it for transactions; writable transactions use BEGIN IMMEDIATE.
    3. Sqlite.ts already enables WAL, sets busy_timeout=5000, foreign keys, and a 32 MiB journal_size_limit. It does not explicitly set synchronous. I did not query the server connection's live effective synchronous mode; it must not be claimed as independently measured here. WAL is already enabled, and a lock busy timeout does not bound filesystem flush latency.
    4. ProviderEventIngestor.ingestNormalized forwards normalized provider events to the event sink. EventSink.write appends events, applies projections, and enqueues effects through commitThenPublish's transaction. The captured worst writes each had one event; this report does not identify their provider or payload type.

    Together, the source and aligned traces strongly support synchronous persistence exposing the main event loop to storage stalls. They do not isolate whether the slow native operation was a COMMIT flush, checkpoint, page read, or another operation inside the transaction.

    Fix areas and useful verification

    • Isolate synchronous database work, including writes and checkpoints, from the HTTP/WebSocket event loop with a bounded worker queue and explicit backpressure. Requests needing persisted state may still wait; event-loop scheduling and keepalives should remain responsive.
    • Measure queue wait separately from native statement, transaction commit, and checkpoint duration. Include effective connection PRAGMAs and redacted event type/count/size so elapsed SELECT spans are not mistaken for query execution time.
    • Review provider/tool update coalescing and transaction batching. The existing traces include ThreadLiveEventCoalescer activity, so this report does not claim coalescing is universally absent. No event-type histogram or T3-specific physical-write amplification measurement was collected.
    • Evaluate WAL synchronous=NORMAL only as an explicit durability decision. It can reduce commit sync work but does not remove synchronous reads or all checkpoint I/O. SQLite documents that recent committed transactions can be lost after power loss with NORMAL. See SQLite WAL and synchronous.
    • Add a controlled regression test with slow database I/O while agents produce updates: assert health/keepalive scheduling, bounded pending writes, preserved ordering and final content, and persistence/error recovery. Reproduce on a disposable disk-backed database, not by filling a user's daily-driver filesystem.

    #14701 covers the related read path. PR #14703 is still open as checked on October 9 and proposes moving selected reads to a worker; it would not by itself isolate these writes.

    Diagnostic scope and limitations

    Diagnostics used t3 trace summary, t3 connect status, t3 service status, retained server traces and boot logs, live connector status, process wait channels, PSI/vmstat, VM XML, ZFS and SMART checks. t3 connect status describes saved setup, not live health. The PATH CLI was older (0.0.44-nightly.20260929.2456) than the running service; the later trace summary was run with the exact installed 2861 executable.

    No database maintenance, PRAGMA change, T3 restart/update, VM storage-path change, or forced stop of another workload was performed by this investigation. A direct cache-pool image path is only a proposed infrastructure mitigation, not an A/B result. No destructive stress reproduction or CI was run. No failure screenshot is attached; the report is based on recorded timing and process evidence.

    Apart from the requested first name in the attribution notice, identifying details are omitted: no personal names, usernames, email addresses, private repository names, environment IDs, thread/project IDs, hostnames, public or private network addresses, hardware serial numbers, absolute home paths, credentials, message contents, or private query parameters. The loopback address, generic application paths, hardware model, software versions, and diagnostic timestamps remain because they describe the failure without naming the user or machine. The raw database and unredacted traces are not attached. A local redacted trace summary has been retained for follow-up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions