Repository navigation
Transient "database is locked" during an event write fails the run but leaves the Claude process running unsupervised #15542
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Thanks for the careful diagnosis and the timeline, @areidyOTH. I confirmed this on
main(40596072), which still has the code from your nightly (73799330). There are two separate bugs here, and it isn't a duplicate of #15447, #5099, or #6097.What I found
The lock fails immediately instead of waiting.
packages/shared/src/nodeSqliteClient.tsnever setsbeginTransaction, so transactions open with a plain deferredBEGIN. An event write reads first (to allocate the turn-item position) and then writes. When a deferred transaction tries to upgrade from its read lock to a write lock while another connection is writing, SQLite returnsSQLITE_BUSYright away and skips the busy handler, sobusy_timeoutnever applies.- Tested with
node:sqliteandPRAGMA busy_timeout = 5000: the deferred upgrade failed in 0 ms withdatabase is locked, whileBEGIN IMMEDIATEwaited about 5 s as expected. - Your 458 ms
EventSink.writespan ending in a rollback matches the immediate failure. A timeout expiry would take at least 5 s. - Upstream
@effect/sql-sqlite-nodeusesBEGIN IMMEDIATEon writable connections. The local client doesn't. - [Bug]: SQLite busy_timeout is never set — concurrent CLI/server writes fail with "database is locked" #5099 only added
busy_timeout. The other writer here was a separate process (t3 auth session issue/revoke), which is exactly what the timeout should wait out.
A failed write ends the run but not the Claude process. In
RunExecutionService.startRootRun, a failed ingest writes a terminal failure throughmakeProviderFailure({ class: "unknown" }), which is the "Provider error / Provider turn failed" item you saw. The fiber then just closes its event subscription and never callsinterruptTurn.ClaudeAdapterV2.interruptTurnis what callsquery.interruptandquery.close, and nothing on this path calls it, sostartTurnkeeps running.- Once the run is
failedwith no running provider turn,run.interruptreturns "Run … is not interruptible," so Stop can't reach the process either. - The process could keep calling tools because the thread was in
full-accessmode (bypassPermissions).
#15447 has the same result for Cursor (run marked failed, provider keeps going) with a different trigger.
Likely fix area
- Transactions: one option is starting event-write transactions with
BEGIN IMMEDIATEso a brief lock waits onbusy_timeout. Retrying the write is another. - Orphaned provider: when ingest fails, the run could interrupt and close the live provider session. That would probably cover [Bug]: Cursor run continues after T3 marks parent and native subagents failed #15447 too.
A maintainer will decide on the fix direction.
- Tested with
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 4, 2026 Another occurrence on the same nightly, plus a variant where the run never leaves
runningSame build as this report (
0.0.46-nightly.20261004.2644,7379933), headlesst3 serveunder a systemd user service on Linux x86_64 (kernel 7.0.0-22), Claude Agent SDK provider (Claude Code 2.1.288) plus Codex. The other writer is the same too: a local MCP bridge runst3 auth session issue/revokearound every API call it makes.Frequency.
orchestration V2 provider event ingestion failedwas logged 8 times between 08:07 and 09:31 UTC on 2026-10-04, across 8 different runs (Claude and Codex, root threads and delegated tasks). In 7 of the 8, a CLIauth session issue/revokewrite had landed 11–1343 ms earlier (6 of them within 600 ms). For the remaining one there was no CLI auth write in the preceding 3 s, so the other writer there is unknown.Trace. 5 of the 8 were still in the rotated trace files, and all 5 fail the same way, fast:
orchestrationV2.EventSink.write(fromproviderTurnStart.start) →sql.transaction→sql.executeofINSERT INTO orchestration_v2_turn_item_positions (thread_id, turn_item_id, ordinal) SELECT ?, ?, COALESCE(MAX(ordinal), ?) + 1 FROM …→TurnItemPositionStoreError←SqlError←Error: database is locked.- The failing statement returned after 0.2–1.4 ms (the whole
EventSink.writespan took 2.8–4.1 ms). That is the immediate busy return described in the triage, not a 5 sbusy_timeoutexpiry.
Variant: the failed terminal is lost as well. In 4 of those 5, the compensating
orchestrationV2.EventSink.writeIfRunCurrentran 5–26 ms later and failed with the samedatabase is locked. Those runs never got a terminal at all, so they stayedrunninginstead of showing "Provider turn failed". An interrupt request on one of them did not settle it either. Only 1 of the 5 (08:52:40Z) got its failed terminal written. This is the path #14856 addresses.In one of the stuck cases (a root Claude thread, failure at 09:20:51Z):
- The Claude CLI kept working and finished the turn normally about 6.5 minutes later (final answer and Stop hook at 09:27:33Z in its transcript). None of it reached the projection, and the UI stayed on "Thinking".
- The user's next message (09:29:05Z) was turned into a steer of the stuck run. The
provider-turn.steereffect failed 5 times with… is not the active turn. A delegated-task completion notice at 09:34:57Z also failed to steer.
In another (a root Claude thread, failure at 09:31:06Z), the provider session was still alive: a later steer was delivered and the agent kept working, but nothing it produced after 09:31:06Z was shown.
Recovery. Only
t3 service restartcleared them. Startup reconciliation cancelled the stuck runs, and with them the in-flight delegated children, whose parents were then woken by the cancellation notices.Happy to re-test once a nightly includes #15488. Since the failed terminal can be lost in the same lock window, the retry and interrupt handling in #14856 still seems needed on top of that.
What happened
A Claude thread failed about 80 s into its first turn with "Provider error: Provider turn failed." Claude itself had not failed: the underlying
claudeprocess kept working with no one watching it. It ran tools (with permission checks bypassed) for another 14+ minutes, and it set up a background systemd timer that competed with a replacement thread the user had started. The UI showed the thread as failed the whole time, and nothing in T3 could stop the process. In the end it had to be found and paused by hand (SIGSTOP).Diagnosis
orchestrationV2.EventSink.write(called fromorchestrationV2.providerTurnStart.start) for the thread failed after 458 ms withTurnItemPositionStoreError←SqlError←Error: database is locked. The transaction was rolled back. The run was then marked failed at 06:59:28.314 with the item "Provider error / Provider turn failed."busy_timeout = 5000is set (apps/server/src/persistence/Layers/Sqlite.ts:19), yet the write failed after about 0.46 s. That pattern suggests SQLITE_BUSY coming back immediately when a deferred transaction tries to upgrade to a write lock while another connection holds it (the busy handler isn't used in that case). This is inference; not verified. Contention was high at that moment: seven thread launches with worktree provisioning were running, and an external helper was issuing and revoking sessions throught3 auth session issue/revokeCLI processes, which write to the same database.providerSessionIdkeeps receivingassistant/tool_use/tool_resultevents until 07:13 UTC (14 min after the failure), and theclaude --output-format stream-json …process was still alive afterwards.Steps to reproduce
t3 auth session issue --ttl 5m --json/t3 auth session revokein a loop from another process.database is locked, the thread changes to "Provider turn failed", whilepsstill shows theclaudeprocess running and the provider events log keeps growing.Version
0.0.46-nightly.20261004.2644 (desktop AppImage, commit 7379933)
Environment
Linux x64 7.0.0-38-generic, Node v26.8.2, claude 2.1.289 (Claude Agent SDK provider, full-access runtime mode)
Evidence
Related issues
#15447 (a Cursor run keeps going after T3 marks it failed): the same "failed run, provider keeps running" problem, seen here with Claude and triggered by a database error rather than a subagent failure. #5099 (busy_timeout never set) was fixed by setting
busy_timeout, but the immediate lock failure under load still happens. #6097 (second backend sharing the database) does not apply here; only one backend process hadstatev2.sqliteopen.Fix applied or workaround
None in T3. The user's replacement coordinator paused the orphaned
claudeprocess withSIGSTOP; it is left in place for inspection.Filed by
Claude Code (claude-opus-5-5) via
t3 triage