Skip to content

[Bug]: T3 Connect discards the failed provisioning stage from client errors #14580

Description

@ECuteri

T3 Connect drops the failed provisioning stage before the error reaches clients. Distinct allocation, Cloudflare tunnel, DNS, and token failures all appear as managed_endpoint_provisioning_failed, so users cannot include the stage in a report or distinguish these failures.

Related incident: #14070. Its maintainer triage identifies the stage as the information needed to investigate. This report isolates the diagnostic information loss; it does not establish the cause of the hosted provisioning incident.

Reproduction

  1. T3 Code desktop 0.0.44 on macOS, Settings → Connections → enable T3 Connect for the local Mac.
  2. Provisioning fails, the toggle returns to off, and the desktop reports:
https://relay.t3.codes/v1/client/environment-links failed:
Relay cannot provision the managed endpoint (managed_endpoint_provisioning_failed).

Real traces from 2026-10-01:

  • Original: f8f444f3bf4f07fc4ff5aacb9243fcc9
  • Fresh retry: 994fe515cbc487d7cf90d24637988a3c

The local cloud.getRelayClientStatus and environment.cloud.makeLinkProof spans succeeded. The relay health route responded successfully, but provisioning still failed. No environment ID reset, credential deletion, or application restart was performed.

Source-level defect

At upstream 5cc99e1c23980d7995a13c47f969b47cb68ed1be, ManagedEndpointProvisioningFailed carries a bounded stage. The linkEnvironment handler in infra/relay/src/http/Api.ts discards it when constructing RelayEnvironmentLinkUnavailableError. The shared contract and web/mobile error presenter consequently cannot expose it.

Expected behavior and focused fix

Preserve the stage as optional diagnostic information in the 503 response and shared client message, alongside the existing reason and trace ID. Keep provider causes, credentials, and resource identifiers out of the response. Retain decoding and message compatibility with older relays that omit the field.

A local regression test reproduces the omission through the actual HTTP API handler with an injected provisioning failure. A proposed patch preserves the stage through HTTP serialization, contract decoding, and shared client presentation; 43 focused tests, targeted lint, formatting, and the three affected package typechecks pass on Node 24.13.1.

The production trace stage and underlying cause still need inspection by a relay operator. A diagnostics patch alone does not restore managed endpoint provisioning.

Activity

  1. iboT2012 commented on Oct 1, 2026

    @iboT2012

    macOS, same managed_endpoint_provisioning_failed when enabling T3 Connect.
    Trace ID: 02b316326f03debe54c241b0c9258cc6

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions