Skip to content

Transient "database is locked" during an event write fails the run but leaves the Claude process running unsupervised #15542

Description

@areidyOTH

What happened

A Claude thread failed about 80 s into its first turn with "Provider error: Provider turn failed." Claude itself had not failed: the underlying claude process kept working with no one watching it. It ran tools (with permission checks bypassed) for another 14+ minutes, and it set up a background systemd timer that competed with a replacement thread the user had started. The UI showed the thread as failed the whole time, and nothing in T3 could stop the process. In the end it had to be found and paused by hand (SIGSTOP).

Diagnosis

  • At 06:59:27.848 UTC, orchestrationV2.EventSink.write (called from orchestrationV2.providerTurnStart.start) for the thread failed after 458 ms with TurnItemPositionStoreError ← SqlError ← Error: database is locked. The transaction was rolled back. The run was then marked failed at 06:59:28.314 with the item "Provider error / Provider turn failed."
  • busy_timeout = 5000 is set (apps/server/src/persistence/Layers/Sqlite.ts:19), yet the write failed after about 0.46 s. That pattern suggests SQLITE_BUSY coming back immediately when a deferred transaction tries to upgrade to a write lock while another connection holds it (the busy handler isn't used in that case). This is inference; not verified. Contention was high at that moment: seven thread launches with worktree provisioning were running, and an external helper was issuing and revoking sessions through t3 auth session issue/revoke CLI processes, which write to the same database.
  • A brief lock while saving one event failed the whole run, instead of the write being retried.
  • When the run failed, the provider session was not stopped. The provider event log for the same providerSessionId keeps receiving assistant / tool_use / tool_result events until 07:13 UTC (14 min after the failure), and the claude --output-format stream-json … process was still alive afterwards.

Steps to reproduce

  1. Start a Claude thread with a long first turn (many tool calls).
  2. While it runs, put write pressure on the database: launch several threads at once, and/or run t3 auth session issue --ttl 5m --json / t3 auth session revoke in a loop from another process.
  3. Once an event write hits database is locked, the thread changes to "Provider turn failed", while ps still shows the claude process running and the provider events log keeps growing.

Version

0.0.46-nightly.20261004.2644 (desktop AppImage, commit 7379933)

Environment

Linux x64 7.0.0-38-generic, Node v26.8.2, claude 2.1.289 (Claude Agent SDK provider, full-access runtime mode)

Evidence

# server.trace.ndjson, span orchestrationV2.EventSink.write, durationMs 458, thread_id <coordinator>
TurnItemPositionStoreError:
    at orchestrationV2.EventSink.write (binCli-*.mjs:83891:21)
    at orchestrationV2.providerTurnStart.start (binCli-*.mjs:229232:59)
    at ServerRuntimeStartup.startEffectWorkerWithRelay (binCli-*.mjs:239491:79)
  [cause]: effect/sql/SqlError: Failed to execute statement
    [cause]: effect/sql/SqlError/UnknownError: Failed to execute statement
        at classifySqliteError (binCli-*.mjs:57423:9)
      [cause]: Error: database is locked
# sibling span sql.transaction -> event db.transaction.rollback, db.name=statev2.sqlite

# thread snapshot
recentRuns: [{status: "failed", startedAt: 06:58:04.087Z, completedAt: 06:59:28.314Z}]
item: {type: "error", title: "Provider error", text: "Provider turn failed."}

# provider/events.<thread>.log, same providerSessionId, after the failure
[06:59:28.449Z] assistant tool_use Bash ...
[06:59:46.157Z] system thinking_tokens ...
...
[07:13:11.038Z] (last event; 3,539 lines total)
$ ps -o stat,etime,args -p <pid>
Tl 24:54 claude --output-format stream-json --verbose --input-format stream-json ... --model claude-opus-5-5[1m] ...

Related issues

#15447 (a Cursor run keeps going after T3 marks it failed): the same "failed run, provider keeps running" problem, seen here with Claude and triggered by a database error rather than a subagent failure. #5099 (busy_timeout never set) was fixed by setting busy_timeout, but the immediate lock failure under load still happens. #6097 (second backend sharing the database) does not apply here; only one backend process had statev2.sqlite open.

Fix applied or workaround

None in T3. The user's replacement coordinator paused the orphaned claude process with SIGSTOP; it is left in place for inspection.

Filed by

Claude Code (claude-opus-5-5) via t3 triage

Activity

  1. juliusmarminge commented on Oct 4, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the careful diagnosis and the timeline, @areidyOTH. I confirmed this on main (40596072), which still has the code from your nightly (73799330). There are two separate bugs here, and it isn't a duplicate of #15447, #5099, or #6097.

    What I found

    The lock fails immediately instead of waiting. packages/shared/src/nodeSqliteClient.ts never sets beginTransaction, so transactions open with a plain deferred BEGIN. An event write reads first (to allocate the turn-item position) and then writes. When a deferred transaction tries to upgrade from its read lock to a write lock while another connection is writing, SQLite returns SQLITE_BUSY right away and skips the busy handler, so busy_timeout never applies.

    • Tested with node:sqlite and PRAGMA busy_timeout = 5000: the deferred upgrade failed in 0 ms with database is locked, while BEGIN IMMEDIATE waited about 5 s as expected.
    • Your 458 ms EventSink.write span ending in a rollback matches the immediate failure. A timeout expiry would take at least 5 s.
    • Upstream @effect/sql-sqlite-node uses BEGIN IMMEDIATE on writable connections. The local client doesn't.
    • [Bug]: SQLite busy_timeout is never set — concurrent CLI/server writes fail with "database is locked" #5099 only added busy_timeout. The other writer here was a separate process (t3 auth session issue / revoke), which is exactly what the timeout should wait out.

    A failed write ends the run but not the Claude process. In RunExecutionService.startRootRun, a failed ingest writes a terminal failure through makeProviderFailure({ class: "unknown" }), which is the "Provider error / Provider turn failed" item you saw. The fiber then just closes its event subscription and never calls interruptTurn.

    • ClaudeAdapterV2.interruptTurn is what calls query.interrupt and query.close, and nothing on this path calls it, so startTurn keeps running.
    • Once the run is failed with no running provider turn, run.interrupt returns "Run … is not interruptible," so Stop can't reach the process either.
    • The process could keep calling tools because the thread was in full-access mode (bypassPermissions).

    #15447 has the same result for Cursor (run marked failed, provider keeps going) with a different trigger.

    Likely fix area

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 4, 2026
  3. whatever1111 commented on Oct 4, 2026

    @whatever1111

    Another occurrence on the same nightly, plus a variant where the run never leaves running

    Same build as this report (0.0.46-nightly.20261004.2644, 7379933), headless t3 serve under a systemd user service on Linux x86_64 (kernel 7.0.0-22), Claude Agent SDK provider (Claude Code 2.1.288) plus Codex. The other writer is the same too: a local MCP bridge runs t3 auth session issue / revoke around every API call it makes.

    Frequency. orchestration V2 provider event ingestion failed was logged 8 times between 08:07 and 09:31 UTC on 2026-10-04, across 8 different runs (Claude and Codex, root threads and delegated tasks). In 7 of the 8, a CLI auth session issue/revoke write had landed 11–1343 ms earlier (6 of them within 600 ms). For the remaining one there was no CLI auth write in the preceding 3 s, so the other writer there is unknown.

    Trace. 5 of the 8 were still in the rotated trace files, and all 5 fail the same way, fast:

    • orchestrationV2.EventSink.write (from providerTurnStart.start) → sql.transaction → sql.execute of INSERT INTO orchestration_v2_turn_item_positions (thread_id, turn_item_id, ordinal) SELECT ?, ?, COALESCE(MAX(ordinal), ?) + 1 FROM … → TurnItemPositionStoreError ← SqlError ← Error: database is locked.
    • The failing statement returned after 0.2–1.4 ms (the whole EventSink.write span took 2.8–4.1 ms). That is the immediate busy return described in the triage, not a 5 s busy_timeout expiry.

    Variant: the failed terminal is lost as well. In 4 of those 5, the compensating orchestrationV2.EventSink.writeIfRunCurrent ran 5–26 ms later and failed with the same database is locked. Those runs never got a terminal at all, so they stayed running instead of showing "Provider turn failed". An interrupt request on one of them did not settle it either. Only 1 of the 5 (08:52:40Z) got its failed terminal written. This is the path #14856 addresses.

    In one of the stuck cases (a root Claude thread, failure at 09:20:51Z):

    • The Claude CLI kept working and finished the turn normally about 6.5 minutes later (final answer and Stop hook at 09:27:33Z in its transcript). None of it reached the projection, and the UI stayed on "Thinking".
    • The user's next message (09:29:05Z) was turned into a steer of the stuck run. The provider-turn.steer effect failed 5 times with … is not the active turn. A delegated-task completion notice at 09:34:57Z also failed to steer.

    In another (a root Claude thread, failure at 09:31:06Z), the provider session was still alive: a later steer was delivered and the agent kept working, but nothing it produced after 09:31:06Z was shown.

    Recovery. Only t3 service restart cleared them. Startup reconciliation cancelled the stuck runs, and with them the in-flight delegated children, whose parents were then woken by the cancellation notices.

    Happy to re-test once a nightly includes #15488. Since the failed terminal can be lost in the same lock window, the retry and interrupt handling in #14856 still seems needed on top of that.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions