Repository navigation
[Bug]: Orchestrator V2 event writes stall the server on a high-latency disk (156 ms fsync), environment flips Offline #15707
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Thanks for the fsync numbers, trace spans, and event counts, @deepontology. That made this very concrete.
What I found
I checked
v0.0.46-nightly.20261004.2644and currentmain(4ee6bfd50e). The write path is the same in both. The only SQLite change since your nightly is #15488, which switches transactions toBEGIN IMMEDIATEbut still commits on the same connection.apps/server/src/persistence/Layers/Sqlite.tsopens a singlenode:sqliteDatabaseSyncin WAL mode. It never setsPRAGMA synchronous, so SQLite runs atFULLand fsyncs the WAL on every commit.packages/shared/src/nodeSqliteClient.tsputs that connection behind a one-permit semaphore and runs statements and commits on the main thread. A slow fsync therefore blocks the event loop and every other statement.- Each provider event is committed on its own.
ProviderEventIngestor.ingestNormalizedcallsEventSink.write, which appends the event, applies projections, and commits in one transaction. At about 156 ms per fsync, the disk can do roughly six commits a second. Your stretches at about nine events a second line up with the I/O wait you saw. - Assistant text is already coalesced in the default
paragraphstreaming mode. Two paths still write on every update:- OpenCode emits a
node.updatedfor each reasoning delta, and those are not coalesced like Claude reasoning and Codex text. - Tool updates (
node.updatedandturn_item.updated) have no streaming filter.
- OpenCode emits a
- That fits a single node ID appearing hundreds of times in a short window. After an abort, events that are already queued still commit one at a time until the queue empties.
- The client errors follow from the stall. The environment descriptor fetch has a 10 s timeout, and the websocket closes after missed pongs. Each reconnect cycles
SessionStore.markDisconnected/markConnectedand reloads the snapshot.
Likely fix area
Some options:
- Use
PRAGMA synchronous=NORMALin WAL mode, which syncs at checkpoint instead of on every commit. - Coalesce OpenCode reasoning and tool updates, and batch queued provider events into one transaction.
- Move the writer off the event loop so a slow commit can't starve keepalives.
#14701 is the read-path side of the same single connection. Open PR #14703 moves that read to a worker but leaves writes on the main thread, so it wouldn't cover this case. A trace slice from a stall would still help show the mix of events.
A maintainer will decide on the fix direction.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 4, 2026 Trace data from a second stall, as requested.
Setup: same machine and build as the report. Two opencode runs started at 16:32 UTC, and the environment flapped again while they streamed. The trace files cover 16:38 to 17:13 UTC.
Event mix, 16:32 to 17:13 UTC, from orchestration_events:
event type count node.updated 18,171 turn-item.updated 2,847 provider-thread.updated 490 message.updated 163 provider-turn.updated 52 node.updated by kind: reasoning 16,549 (91.1%), tool_call 1,448 (8.0%), assistant_message 158 (0.9%), root_turn 12.
By provider: opencode 18,116 (99.7%), claudeAgent 55.
Most-updated native items, all OpenCode reasoning nodes:
- prt_107c694da001... 1,049 updates
- prt_107cdb427001... 715 updates
- prt_107ca6c7a001... 638 updates
- prt_107cf4ab4001... 579 updates
Worst minute, 17:10:00 to 17:11:00 UTC, from server.trace.ndjson:
- orchestrationV2.EventSink.write: 487 spans, 106,295 ms total
- sql.transaction: 514 spans, 108,681 ms total
- 106.3 s of commit time in a 60 s minute is 177% saturation on the single-threaded writer
- durations: p50 100 ms, p95 1,125 ms, max 3,314 ms
- busy time per second stayed mostly above 1,000 ms, so the queue never drained
The pattern holds across every minute of the incident. All 37 minutes in the trace show 104 to 112 s of transaction time per minute.
The first incident showed the same mix: 15:50 to 16:15 UTC had 7,619 node.updated, of which 6,022 (79%) were reasoning and 1,475 (19.4%) were tool_call, with opencode responsible for 7,554 (99.1%).
Artifacts in this gist: https://gist.github.com/deepontology/ac5bdb621948c59bebde99efae4768da
- trace-slice.jsonl: every span in the worst minute, fields are ts, span, ms, status, trace. 4,501 rows. Attributes, IDs, and payloads stripped.
- per-minute.csv: sink and transaction counts and busy time for 16:38 to 17:13 UTC.
- per-second-worst-minute.csv: per-second commit count and busy time for 17:10 UTC.
This matches your triage. OpenCode reasoning deltas dominate, that path has no coalescing, and the commits serialize on the one connection.
Submitted with an opencode agent 🤖
Another data point, on fast storage, focused on how much this writes rather than fsync latency.
Setup:
t3 serve0.0.46-nightly.20261005.2689 as a systemd user service, Debian 13, Crucial P3 NVMe, btrfs (compress=zstd:3), 16 GB RAM. Mostly Claude agents, often 5 to 10 at once, plus Codex delegated tasks.statev2.sqliteis 1.3 GB with 0 free pages. Growth oforchestration_eventsby day (MB ofpayload_json):Day Events MB 09-20 to 10-02, typical 2k to 18k 1 to 15 10-03 (V2 in use from here) 36,883 84 10-04 40,742 151 10-05 39,984 114 By type, all time:
event_type count MB turn-item.updated46,778 250 thread.activity-appended64,191 113 node.updated29,967 31 provider-turn.updated13,084 15 turn-item.updatedstores the whole item on each update, so long tool outputs are written again on every update. Singleorchestration_v2_projection_turn_itemsrows reach 680 KB, and that table is another 200 MB. One long-running thread with its subagents holds 264 MB of events and items. Nothing prunes events, including those of archived threads.Read side (#14701 territory), from the trace: a
message.dispatchtransaction took 6.1 s, of which 3.4 s wasSELECT payload_json FROM orchestration_v2_projection_turn_items WHERE thread_id = ?for a thread of 1,769 items (3.4 MB). The host was under swap pressure, so pages were cold, but the query loads every item of the thread on the main thread. The server logs "event loop stalled" of 2 to 8 s several times a day, once 41 s, and remote clients drop and resync on each.Coalescing tool
turn_item.updatedwrites as you suggested would cut most of this volume. A retention rule for events of archived or old threads would bound the rest; I can open an Ideas discussion for that if it helps.I’m seeing similar Offline/Reconnecting symptoms on Windows desktop with a WSL2 backend, using a newer V2 nightly. I captured the server’s main thread waiting in
fsync()for an ext4 journal commit while T3 was showing Offline/Reconnecting.Environment
Desktop and backend:
0.0.46-nightly.20261005.2702WSL distribution: Ubuntu 24.04
“WSL only” enabled
Local backend:
http://127.0.0.1:3773~/.t3resolves inside the WSL home directoryFilesystem:
/dev/sdf, ext4, mounted at/Mount options:
rw,relatime,discard,errors=remount-ro,data=ordered
Filesystem capacity:
Filesystem Size Used Avail Use% Mounted on /dev/sdf 1007G 386G 570G 41% /Physical drive details and Windows host free space have not been collected. I have not independently verified the live database’s exact path.
Workload and symptoms
I started a coordinator with four ordinary worker threads through T3’s thread-management tools. Multiple Flutter compilation/test processes were running as part of their work.
The environment showed Offline/Reconnecting while Claude and build/test processes remained alive. This is the first occurrence I have observed in my setup.
Process and I/O evidence
At approximately 18:34 UTC+7 on October 6, the T3 server process showed:
PID PPID STAT WCHAN COMMAND 1674 1673 Dsl+ folio_wait_bit_common node-MainThreadSix visible Claude processes had PID
1674as their parent.The four interval samples from
Sample | Blocked processes (b) | I/O wait (wa) | CPU idle (id) -- | -- | -- | -- 1 | 10 | 38% | 53% 2 | 8 | 24% | 61% 3 | 7 | 27% | 67% 4 | 8 | 27% | 59%vmstat 1 5, excluding the first since-boot report, showed:Swap usage was approximately 1.4 GiB. Three Dart frontend processes held approximately 4.65 GiB of resident memory combined.
The server still owned port 3773. Its listening socket’s Recv-Q was 7 at 18:34 and 1 at 18:44.
Kernel stack captured while T3 was still Offline/Reconnecting
At approximately 18:44–18:45 UTC+7:
sudo cat /proc/1674/stack[<0>] jbd2_log_wait_commit+0xb4/0x130
[<0>] jbd2_complete_transaction+0x73/0xb0
[<0>] ext4_fc_commit+0x64b/0x9f0
[<0>] ext4_sync_file+0x1e4/0x2b0
[<0>] vfs_fsync_range+0x49/0x90
[<0>] do_fsync+0x40/0x90
[<0>] __x64_sys_fsync+0x17/0x20
[<0>] x64_sys_call+0x1729/0x20f0
[<0>] do_syscall_64+0x73/0x990
[<0>] entry_SYSCALL_64_after_hwframe+0x76/0x7eThis captures a synchronous filesystem commit wait in the server’s main thread. It does not identify the file being synchronized or measure the duration of that operation, so I cannot yet confirm that SQLite is the caller or that this has the same root cause.
Would specific T3 traces or further storage diagnostics help distinguish the V2 event-write path from WSL storage contention caused by concurrent builds/tests?
Another data point: an NVMe SSD on ext4, where the trigger is other processes writing heavily rather than a slow disk. On SSH-managed environments the stall also ends up killing running agents.
Setup
0.0.46-nightly.20261007.2774, started by the desktop app as an SSH-managed environment (~/.t3/ssh-launch/…/run-t3.sh)- NixOS, kernel 6.18.55, 58 GB RAM, no swap
~/.t3on ext4 (defaultdata=ordered), Crucial P3 4 TB NVMe (CT4000P3SSD8, firmwareP9CR30A)statev2.sqliteis 2.3 GB- Workload: several Claude agents, each building a Rust workspace in its own worktree (
cargo build/cargo nextest,debuginfo=2)
Numbers
- fsync of a 4 KB file in
~/.t3/userdata: 26–45 ms normally, spiking to 5.5 s while the builds were writing about 60 MB/s. - At the worst point there were 8.7 GB of dirty pages and
/proc/pressure/ioshowedfull avg10=81. Memory pressure was 0, so tools like btop showed nothing wrong. server.eventLoop.stallreached 15–16 s several times. CPU use during the stalls was low (utilization0.2–0.5, a few tens of ms of CPU), so the main thread was waiting, not computing.- Trivial statements took seconds:
UPDATE auth_sessions SET last_connected_at = ?took 4.2 s andSELECT * FROM scheduled_tasks WHERE enabled = 1 …took 7.5 s. Onesql.transactiontook 23.6 s. - The only PRAGMAs in the traces are
journal_mode = WAL,busy_timeout = 5000,foreign_keys = ONandjournal_size_limit. That matches triage:synchronousis left atFULL.
With ext4
data=ordered, a journal commit can wait for other files' dirty data to be written out first. So a small WAL fsync gets queued behind the build output, even on an SSD. This probably answers the WSL comment's question: concurrent builds and tests are enough to trigger it without a slow disk.What it does to SSH-managed environments
The stall gets worse than "Offline" here:
- The event loop blocks for 10–16 s.
- The desktop's SSH session drops, the client reconnects, and the launcher starts a new server. That happened 6 times in about 20 minutes.
- Startup recovery then ends whatever was running, e.g.
V2 orchestration recovery completed { terminalizedRuns: 3, stoppedSessions: 4 }at 18:22:51.
Each restart kills in-flight agent turns and their child processes. Inside the agents the builds show up as exit 137, and the agents reported this to me as "out of memory", even though there was no kernel or oomd kill and about 50 GB of RAM was free.
A second
t3 serveon the same~/.t3(the #5749 / #6097 situation) makes this worse. Each slow commit holds the database lock for seconds, and the other server fails withLockTimeoutError: database is lockedoncebusy_timeoutruns out.Mitigations on my side
Fewer concurrent builds, lower
vm.dirty_bytes/vm.dirty_background_bytes, anddebug = "line-tables-only"for dev builds. These shrink the spikes but don't remove them.PRAGMA synchronous=NORMALin WAL mode, or moving the writer off the main thread, as suggested in triage, would fix it at the source.Trace segments covering the stalls and restarts are available if useful.
Investigated and written by Claude Opus 5.5 running in Claude Code (via T3 Code).
- added a commit that references this issue
on Oct 7, 2026 Another data point: the same disconnect cycle on fast storage. My fsync latency is about 30 times lower than in the original report, so slow disks aren't the only trigger.
Setup
0.0.46-nightly.20261005.2702(Arch,t3code-nightly-bin), desktop app, local environment only- Omarchy (Arch), kernel 7.2.5, 32 cores, 125 GB RAM
~/.t3on btrfs (compress=zstd:3) on LUKS, Samsung 990 PRO 2 TB NVMe- fsync of a 4 KB write in
~/.t3/userdata: 5.0 ms average (p50 4.9 ms, max 5.6 ms, 50 samples) - Workload: Claude agents only (
claudeAgent)
Symptoms
The add-project flow showed "Environment unavailable – omarchy is not connected", and the connection kept dropping and coming back. In
server.trace.ndjson, top-levelEnvironmentRpc.subscribespans were cut short again and again between 18:03 and 18:13 UTC, each lasting 5 to 133 seconds (8.4 s, 16 s, 48 s, 93 s, 133 s, 60 s, …). The busiest minute of event writes, 18:11 UTC with 719 events, falls inside that window. The rotated server trace files filled up quickly: 10 MB in about 4 minutes, mostlysql.execute,db.transaction.commit,orchestrationV2.EventSink.writeandThreadSettlementServiceV2.settleThread.The desktop process had been running for 7h40m with a 9.6 GB memory peak (reported by systemd for the app's scope). A full relaunch fixed it for now.
Database
statev2.sqlite: 564 MB, 0 free pages. The legacystate.sqlite(253 MB) is still on disk.- Largest tables:
orchestration_events265 MB,orchestration_v2_projection_turn_items101 MB,projection_thread_activities91 MB orchestration_v2_eventsis empty. Every event is written toorchestration_events.
Events by type, all time:
event_type count MB turn-item.updated12,046 99.5 thread.activity-appended20,567 89.5 node.updated6,720 6.0 message.updated4,166 3.3 provider-turn.updated3,960 3.6 turn-item.updatedaverages about 8.3 KB per event, which looks like the full item rewritten on every update rather than only what changed.Daily growth:
Day Events MB 10-05 639 0.6 10-06 17,528 86.2 10-07 (partial) 7,991 29.2 With fast storage, disk speed doesn't explain the stalls. The large database, plus the read-path blocking from #14701, seems to be enough to block the main thread even at 5 ms per commit.
I have a small patch ready for the first option in the triage comment:
PRAGMA synchronous = NORMALinlayerSetup(apps/server/src/persistence/Sqlite.ts), next to the other pragmas. It is one line plus a test that the pragma is applied.Before opening a PR, can a maintainer confirm this direction? It changes a default: in WAL mode,
NORMALstops fsyncing the WAL on every commit. After an OS crash or power loss the last few committed transactions can be lost. An application crash is safe, and WAL withNORMALcannot corrupt the database. Checkpoints still fsync.A quick measurement on an NVMe disk (Windows,
node:sqlite, WAL, 1000 single-row transactions): 0.431 ms per commit atFULL, 0.050 ms atNORMAL. That is a fast disk, so it does not reproduce the 156 ms case from the report. I can't confirm that this alone stops the Offline flapping, and it does not touch event volume or moving the writer off the main thread.If this direction is fine, I'll open the PR with disk-backed timings. If you would rather go another way, I'll leave it.
Another data point on fast storage (NVMe + btrfs + LUKS), with disk-backed
FULLvsNORMALtimings from the same disk.Setup
0.0.46_nightly.20261006.2735(Arch,t3code-nightly-bin), desktop app, local environment only- Omarchy, kernel 7.2.8, 16 GB RAM
~/.t3on btrfs (compress=zstd:3) on LUKS, XPG GAMMIX S70 BLADE 1 TB NVMe (SMART clean, 4% used, 0 media errors)- Workload: Claude, Codex and OpenCode, with delegated tasks across providers
Stall evidence
A watcher polled
/.well-known/t3/environmentevery 30 s and captured system state when it failed. It recorded 11 stalls, each a 5 s timeout, while the UI showed reconnecting:- The server main thread was in
Dstate with wchanfolio_wait_bitorwait_log_commit. /proc/pressure/ioshowedsome avg10between 17 and 42, while CPU and memory pressure were about 0.- In 8 of the 11 snapshots, the T3 server was the only userspace process in
Dstate.
During one stall (16:50 to 16:53 local time), queries that are normally fast stretched out:
span median max SELECT * FROM scheduled_tasks …0 ms 23.4 s getLimitRecoveryCandidates3 ms 23.4 s orchestrationV2.EventSink.write— 1 to 6 s client ConnectionDriver.connect— 10 s Write volume
- With agents running, the server process wrote 64 MB to disk in 30 s, across 26,598 write syscalls (
/proc/<pid>/io). statev2.sqliteis 2.5 GB.orchestration_eventsholds 1.46 GB in 467k rows.payload_jsonper day: 10-05 39 MB, 10-06 42 MB, 10-07 161 MB, 10-08 230 MB.- 10-08 to 10-09 by type:
turn-item.updated39,535 rows (182 MB),node.updated88,396 rows (141 MB).
NOCOW does not fix it
This confirms the original report. Rewriting
statev2.sqliteas achattr +Cfile cut it from 157,192 extents to 92. Stalls returned within 40 minutes, with the same wchan values.FULLvsNORMALon this diskTest conditions:
node:sqliteon Node v26.11.1, SQLite 3.53.4, WAL, file on the same btrfs/LUKS volume as~/.t3/userdata. Each run is 1000 single-row transactions (BEGIN IMMEDIATE+ insert of a ~4.6 KB payload +COMMIT), roughly oneturn-item.updatedeach. The disk was idle.p50 p99 max total raw 4 KB write + fsync 1.56 ms 2.12 ms (p95) 2.26 ms — synchronous=FULL1.71 ms 3.03 ms 6.66 ms 1.75 s synchronous=NORMAL0.03 ms 0.13 ms 8.80 ms 0.08 s On an idle disk,
FULLalready costs about one fsync per event. Under load the stall snapshots show those waits growing to seconds, and every one happens on the main thread.NORMALremoves the per-commit wait, leaving only checkpoint syncs. That is themaxoutlier in theNORMALrow.@Sypher760-gif, happy to run your patch on this machine and report whether the stalls stop under the same workload.
Note
gpt-6.1-sol · Codex running in T3Code, responding on behalf of Max.
Another occurrence on an NVMe-backed Linux VM, with 104–122 second main-thread stalls during one-event V2 writes, plus further stalls after the competing emulator workload stopped.
This supports the responsiveness failure tracked here. It does not establish that T3 caused the underlying storage slowdown or that the SSD is defective.
Observed behavior and impact
On 2026-10-09 the macOS client showed:
T3 Connect · Reconnecting: Could not connect relay environment.
The server process remained alive, but even local HTTP requests timed out. Remote access recovered without a restart when disk pressure fell, then additional long stalls appeared. This prevented reliable interaction with running work in the remote environment.
Expected: storage contention can delay persistence, but a blocked database operation should not freeze the main HTTP/WebSocket event loop and cause the entire environment to disappear behind a generic relay error.
Environment
- Running server:
0.0.46-nightly.20261009.2861, Linux x64, systemd user service; kernel6.8.0-146-generic. The matching release tag points to3b6af0bd1600f034b0466e1b8ff017fbabeba910. - Remote connection: T3 Connect with managed cloudflared
2026.10.0, accessed from the macOS desktop app. The desktop version was not collected. - Linux VM: 6 vCPUs, 6 GB allocated RAM (5,919 MiB visible), ext4 on LVM; root filesystem about 94% used with about 31 GB available during diagnosis. About 2 GB of swap was occupied during the severe incident. Swap occupancy alone does not establish active swapping as the cause.
- Live V2 database:
~/.t3/userdata/statev2.sqlite, approximately 12 GiB by roundedls -lhat 13:03 UTC; oldstate.sqliteapproximately 7.4 GiB. The V2 WAL was approximately 13 MiB at that snapshot. These are file sizes, not measured event payload totals or growth rates. - Storage route: guest ext4/LVM → virtio raw virtual disk (
cache=writeback,discard=unmap) → Unraid/mnt/userFUSE (fuse.shfs) → ZFS cache pool → Crucial CT500P2SSD8 NVMe SSD. The image physically resides on the cache SSD pool. No explicit QEMU IOThread was configured. - Hypervisor: about 16 GB RAM total; sampled available memory about 3 GB, ZFS ARC about 2 GB. These host measurements were after the worst stall.
Correlated trace evidence
All times below are UTC on 2026-10-09. The first column identifies the end of the correlated activity, not its start. Durations come from T3's trace; stall duration is the event-loop span's
delayMaxMsattribute, not the tiny duration of the monitoring span itself.End time sql.transactionorchestrationV2.EventSink.writeEvent count on write Event-loop delay 12:20:05 104,464.394 ms 104,467.450 ms 1 104,495 ms 12:23:21 122,614.470 ms 122,616.425 ms 1 122,480 ms 12:57:09–10 69,038.483 ms and a separate 44,107.001 ms transaction 44,108.285 ms 1 69,800 ms The later row contains overlapping spans, not one transaction attributed wholly to one event. A
sql.executeSELECT span ending at 12:57:09 lasted 69,060.493 ms. Its elapsed time can include waiting for the shared connection, so this does not establish a 69-second query execution cost.The monitor reported relatively little CPU use over the long sampled intervals:
Stall User CPU System CPU RSS Major page faults 104,495 ms 2,501 ms 2,544 ms 506 MiB 158 122,480 ms 785 ms 727 ms 535 MiB 28 69,800 ms 1,867 ms 1,193 ms 548 MiB 56 These counters cover the monitor's sample interval, not an isolated native SQLite call. High event-loop utilization here must not be read as high CPU utilization.
A redacted scan of retained trace files, collected at 13:06 UTC and covering span endings from 12:18:04 to 13:05:15, contained:
- 23
server.eventLoop.stallrecords. - 7,100
sql.transactionspans. - 4,591
orchestrationV2.EventSink.writespans. - 43,258
sql.executespans. - 9,446
ThreadSettlementServiceV2.settleThreadspans; several reached about 123 seconds during the severe stall.
These are trace span counts, not unique database events, physical write counts, or additive time on disk. The retained window is incomplete before 12:18 due to log rotation. Long-lived subscription spans were excluded from interpreting request latency.
Process, kernel, HTTP and relay evidence
During the severe incident:
- The server's main thread was sampled in
Dstate, then with wait channeljbd2_log_wait_commit, an ext4 journal-commit wait. No complete native stack or syscall-to-file attribution was captured. /proc/pressure/io:some avg10=96.47%,full avg10=87.10%.vmstatinterval samples: 47–61% CPU I/O wait, six blocked tasks; load average 33.14 on six vCPUs.- HTTP on
127.0.0.1:3773timed out after four seconds. The public environment endpoint timed out after five seconds from the VM and fifteen seconds from the Mac. - Cloudflared reported canceled
/api/t3-connect/healthand/api/t3-connect/mint-credentialrequests and QUIC inactivity timeouts. Local HTTP failure means this was not solely a relay/network outage. - No relevant kernel warning was found in the checked 20-minute window.
Around 12:24:24, without restarting or changing configuration, local HTTP returned 200 in 83 ms and the public endpoint returned 200 in 163 ms. I/O PSI fell to
some avg10=1.10%,full avg10=0.87%; the connector then reported connected/fresh. This was temporary recovery, not a demonstrated fix.Competing workload and storage health
An Android emulator started at 12:11:47, used about 2.3 GB RSS, and
/proc/<pid>/ioshowed about 17.1 GB written after roughly twelve minutes. Its QA workload deliberately tested Android guest storage exhaustion using approximately 5.6 GB filler writes and flushes.The fault was confined to the emulator's
/datavirtual filesystem. The probe retained a 24 GiB reserve on the Linux VM filesystem. The final probe recorded 118,784 bytes free inside Android before the test, then removed the filler and restored approximately 5.6 GB free. This was not evidence that T3's filesystem ran out of space. The workload is strongly correlated with the early incident but was not isolated in a controlled A/B test.The emulator was absent by the 12:40 check and again at the follow-up process check. The 69.8-second stall at 12:57 therefore shows that stopping it did not eliminate all pauses. Remaining workload, delayed writeback, database read cost, memory pressure, and storage queueing have not been separated.
After recovery, SSD SMART showed PASSED, critical warning 0, zero media/data-integrity errors, 46°C and 5% used. ZFS was ONLINE with no known data errors; the pool had about 154 GB free. Historical ZFS averages showed approximately 237 ms total write wait versus 3 ms device write wait and 604 ms synchronous queue wait, but these were not measurements from the exact outage. Idle follow-up samples were about 2 ms total write wait and very low NVMe utilization. This suggests investigating queueing through the storage stack; it does not prove FUSE caused the stall or rule out SSD latency under load.
Relevant code in this exact release
I checked the matching release source, rather than relying only on current main:
nodeSqliteClient.ts: connection construction opensNodeSqlite.DatabaseSync. Statement execution directly invokesstatement.all()/statement.run()inside the Effect runtime. An Effect wrapper does not move these native synchronous calls to another thread.- The same adapter uses a one-permit semaphore around one connection and keeps it for transactions; writable transactions use
BEGIN IMMEDIATE. Sqlite.tsalready enables WAL, setsbusy_timeout=5000, foreign keys, and a 32 MiBjournal_size_limit. It does not explicitly setsynchronous. I did not query the server connection's live effective synchronous mode; it must not be claimed as independently measured here. WAL is already enabled, and a lock busy timeout does not bound filesystem flush latency.ProviderEventIngestor.ingestNormalizedforwards normalized provider events to the event sink.EventSink.writeappends events, applies projections, and enqueues effects throughcommitThenPublish's transaction. The captured worst writes each had one event; this report does not identify their provider or payload type.
Together, the source and aligned traces strongly support synchronous persistence exposing the main event loop to storage stalls. They do not isolate whether the slow native operation was a COMMIT flush, checkpoint, page read, or another operation inside the transaction.
Fix areas and useful verification
- Isolate synchronous database work, including writes and checkpoints, from the HTTP/WebSocket event loop with a bounded worker queue and explicit backpressure. Requests needing persisted state may still wait; event-loop scheduling and keepalives should remain responsive.
- Measure queue wait separately from native statement, transaction commit, and checkpoint duration. Include effective connection PRAGMAs and redacted event type/count/size so elapsed SELECT spans are not mistaken for query execution time.
- Review provider/tool update coalescing and transaction batching. The existing traces include
ThreadLiveEventCoalesceractivity, so this report does not claim coalescing is universally absent. No event-type histogram or T3-specific physical-write amplification measurement was collected. - Evaluate WAL
synchronous=NORMALonly as an explicit durability decision. It can reduce commit sync work but does not remove synchronous reads or all checkpoint I/O. SQLite documents that recent committed transactions can be lost after power loss with NORMAL. See SQLite WAL and synchronous. - Add a controlled regression test with slow database I/O while agents produce updates: assert health/keepalive scheduling, bounded pending writes, preserved ordering and final content, and persistence/error recovery. Reproduce on a disposable disk-backed database, not by filling a user's daily-driver filesystem.
#14701 covers the related read path. PR #14703 is still open as checked on October 9 and proposes moving selected reads to a worker; it would not by itself isolate these writes.
Diagnostic scope and limitations
Diagnostics used
t3 trace summary,t3 connect status,t3 service status, retained server traces and boot logs, live connector status, process wait channels, PSI/vmstat, VM XML, ZFS and SMART checks.t3 connect statusdescribes saved setup, not live health. The PATH CLI was older (0.0.44-nightly.20260929.2456) than the running service; the later trace summary was run with the exact installed2861executable.No database maintenance, PRAGMA change, T3 restart/update, VM storage-path change, or forced stop of another workload was performed by this investigation. A direct cache-pool image path is only a proposed infrastructure mitigation, not an A/B result. No destructive stress reproduction or CI was run. No failure screenshot is attached; the report is based on recorded timing and process evidence.
Apart from the requested first name in the attribution notice, identifying details are omitted: no personal names, usernames, email addresses, private repository names, environment IDs, thread/project IDs, hostnames, public or private network addresses, hardware serial numbers, absolute home paths, credentials, message contents, or private query parameters. The loopback address, generic application paths, hardware model, software versions, and diagnostic timestamps remain because they describe the failure without naming the user or machine. The raw database and unredacted traces are not attached. A local redacted trace summary has been retained for follow-up.
- Running server:
Before submitting
Area
apps/server
Steps to reproduce
Expected behavior
The server keeps answering keepalives while agents stream, or degrades without dropping clients.
Actual behavior
Orchestrator V2 (the OrchestrationV2 migration; tables prefixed orchestration_v2_) persists each provider stream delta as an event. The server process sat in uninterruptible I/O wait for long stretches. Sampled wchan values: btrfs_btree_wait_writeback_range, write_all_supers, folio_wait_bit_common, wait_log_commit. SQLite transactions reached 1.98 s. Trace segments were dominated by orchestrationV2.EventSink.write spans of 40 to 500 ms each. Clients reported:
The server log shows repeated SessionStore.markConnected and markDisconnected cycles. The Android client stayed logged in and showed the same offline state.
Event volume: 10,835 events were appended between 15:50 and 16:15 UTC, about 5,500 in the worst ten minutes, so roughly nine writes per second. In the last 5,000 rows there were 416 distinct node IDs, and one node ID appeared 405 times within 2,000 rows. This matches the mechanism in #14701: one synchronous SQLite connection on the main thread, so a slow commit blocks every other request, and each reconnect repeats the work.
Earlier builds with several concurrent runs did not show this on the same machine.
Impact
Major degradation or frequent failure. T3 Code is unusable while agents run on this machine.
Version or commit
0.0.46-nightly.20261004.2644 (Arch package t3code-nightly-bin 0.0.46_nightly.20261004.2644-1)
Environment
Arch Linux, Wayland. Root and /home on crypto_LUKS + btrfs over a Toshiba HDWD110 HDD, compress=zstd:3. Desktop and Android clients on the same account.
What I tried
Workaround
Reverting to the previous nightly, 0.0.45_nightly.20260930.2468, which predates the Orchestrator V2 event store. state.sqlite is intact, so this restores normal operation on this hardware. I will update after running it for a while.
Open questions
Logs or stack traces
Happy to attach server trace segments, the wchan samples, and the fsync benchmark script if useful.
Related
#14701, #15609, #15442, #14871
Submitted with an opencode agent 🤖