Skip to content

feat(threads): opt in to automatic resume after connection loss #13740

Description

@coygeek

Summary

Add an opt-in setting that continues a task in the same conversation after a confirmed network interruption stops it. Keep chat unchanged during healthy work, and show a compact interruption notice only while needed. Initial automatic recovery supports native Codex using its built-in OpenAI connection.

Problem to solve

A temporary network interruption can stop an active provider turn. Restoring the client connection does not itself continue that task, so the user must notice the interruption and send a continuation manually. Users need an optional way to resume eligible interrupted work without duplicating a task that kept running or finished while the client was disconnected.

Proposed behavior

  • Add "Resume after connection loss" beside "Continue threads after restarts" in General settings, and beside the restart control in mobile Maintenance settings. Default it to off for fresh and existing installations. Use the helper text "Automatically resume tasks interrupted by a lost connection when connectivity returns."
  • Persist the preference for the selected environment and keep it independent of restart continuation. Disclose that initial support is native Codex with the built-in OpenAI connection. Custom endpoints, proxies, and other providers retain their existing retry behavior.
  • Reconcile the task when a client reconnects. Catch up with running or completed work without starting another turn. A quiet task or a lost client connection alone is not evidence that the provider stopped.
  • When a supported provider confirms that a connection failure ended its turn, wait for the required provider connection to return, then continue in the same conversation with its existing context and work. Do not replay the original prompt or compete with the provider's own retries.
  • Show "Connection lost. Reconnecting…" while the client cannot verify the task state. Show "Connection lost. Waiting to resume…" only for a confirmed interruption with recovery enabled. Show "Task resumed" briefly after matching provider progress confirms recovery, then remove it.
  • Let Stop, disabling the setting, newer work, completion, and pending approvals or questions take precedence. If automatic recovery is unsupported or fails, show a concise explanation and preserve manual continuation.

Acceptance criteria

  1. The preference defaults to off, persists per environment, and changes independently of restart continuation. When off, the application initiates no recovery turns.
  2. With native Codex on its supported built-in OpenAI connection, a confirmed network interruption can recover in the same conversation once the provider connection is usable. Cover Wi-Fi loss, Ethernet loss, and DNS failure. A restored client socket alone must not authorize continuation.
  3. Client-only disconnection, task completion while disconnected, and long quiet tool operations create no extra turns. Repeated reconnects, multiple clients, and provider retries cannot create overlapping or duplicate continuations.
  4. Recovery waits with backoff while connectivity is unavailable and does not repeatedly create failed turns. Authentication failures, quota limits, and unrelated task errors do not qualify for automatic recovery.
  5. Stop, disabling the preference, newer work, completion, and pending user interaction prevent pending recovery from overriding user intent. Recovery must not bypass approvals or questions.
  6. Desktop, web, and mobile show no new chat controls during healthy work. Interruption notices are compact, accessible, and do not steal focus. A success notice requires confirmed provider progress and disappears shortly afterward.
  7. Settings disclose the initial provider scope. Unsupported confirmed interruptions receive an explanation instead of an implied promise of recovery. Verify supported local and remote client connections without claiming automatic recovery for custom routes or other providers.

Affected area

Environment settings, provider continuation, and thread interruption feedback in web, desktop, and mobile clients.

Non-goals

Persistent execution-status controls, diagnostic panels, activity timestamps, manual status checks, recovery inferred from inactivity, automatic recovery from arbitrary errors, and support for custom provider routes in the initial implementation. Updates, crashes, and machine restarts remain governed by the separate restart setting.

Implementation

Proposed implementation: PR #13744. The PR contains the support boundary, verification results, and before/after images and a simulated-recovery video.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedenhancementRequested improvement or new capability.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions