Skip to content

fix(perps): read order-path data over HTTP during a WebSocket reconnect - #10590

Draft
abretonc7s wants to merge 9 commits into
mainfrom
TAT-4041-order-path-http-reads
Draft

abretonc7s wants to merge 9 commits into
mainfrom
TAT-4041-order-path-http-reads

Conversation

@abretonc7s

@abretonc7s abretonc7s commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Explanation

Stacked on #10589. Draft until a transport stress test confirms that always reading over HTTP is the right long-term policy for the order path. The alternative is a WebSocket-first policy with HTTP fallback.

Pre-order reads depend on the WebSocket. During a reconnect, HyperLiquidClientService clears its WebSocket info and subscription clients and keeps the HTTP exchange and HTTP info clients. getInfoClient() without { useHttp: true } requires the WebSocket clients, so several reads on the order path threw CLIENT_NOT_INITIALIZED:

  • the #ensureReady gate and asset-map rebuild
  • unified-account and referral setup
  • HIP-3 DEX balances and spot metadata
  • the margin-mode lock (open orders, TWAP history, active asset data)
  • the open-order read in updatePositionTPSL

These reads now go over HTTP, like the metadata, price, position and builder-fee reads already did.

getMaxLeverage capped orders at 3x during a reconnect. It required the WebSocket clients before an HTTP metadata read, and its error path fell back to the 3x default. Orders above 3x were then rejected with ORDER_LEVERAGE_INVALID. The WebSocket gate is removed.

Pre-order reads fail on a transient 429. These reads share HyperLiquid's per-IP REST budget with charts and history. They now retry a 429 up to twice with jittered exponential backoff, at most about 1.5 s: metadata, prices, spot metadata, positions and balances, open orders, TWAP history and asset data. Other errors are not retried. Setup reads and exchange writes are not retried.

Trade-offs to settle before this leaves draft:

  • Every order uses these HTTP reads, not only orders placed during a reconnect.
  • Mobile HTTP has its own failure modes: nitro-fetch aborts and Cronet errors (TAT-3871).
  • More of the per-IP REST budget is used, plus retries during a rate limit.

Reproduction

I ran the real headless PerpsController, SDK transports, signer and BTC resting-limit-order path against HyperLiquid testnet. The fault was introduced at the SDK fetch boundary or the WebSocket transport, and each scenario was run on unfixed source and then on 9eb852d, which contains this change.

Controlled condition Unfixed source With this change
One HTTP 429 on pre-order allMids (ecbc4e9) placeOrder returned 429 Too Many Requests. The retry reached testnet and got HTTP 200. Order 61421338052 was observed open, then canceled.
Replacement WebSocket delayed 8 s during provider.reconnect() (merge-base 1f15b7c) An order submitted while connecting returned CLIENT_NOT_INITIALIZED. Order 61421532628 succeeded while the reconnect was still pending. It was observed open, then canceled by ID.

The 429 was a constructed response. These faults did not occur organically, and no Mobile or Extension UI was exercised.

Validation

  • Provider, controller and utils suites: 51 passed, 2669 passed.
  • The TAT-4041 provider tests fail with CLIENT_NOT_INITIALIZED when these source changes are reverted. A per-site check reverted each changed read on its own, and a test caught every one.
  • 9 unit tests for the retry helper. rateLimitRetry is not exported from the package.

References

Checklist

  • I've updated the test suite for new or updated code as appropriate
  • I've updated documentation (JSDoc, Markdown, etc.) for new or updated code as appropriate
  • I've communicated my changes to consumers by updating changelogs for packages I've changed
  • I've introduced breaking changes in this PR and have prepared draft pull requests for clients and consumer packages to resolve them

Orders submitted while the Perps connection reconnects failed with
CLIENT_NOT_INITIALIZED although the HTTP exchange client stays available.

- Read order-path data over HTTP: margin-mode lock, HIP-3 balances and
  spot metadata, unified-account and referral setup, asset-map rebuilds,
  and the #ensureReady gate. getMaxLeverage no longer requires the
  WebSocket clients, which forced a 3x cap during a reconnect.
- When an action waited on a disconnect or reinitialization, wait up to
  ConnectionTimeoutMs for the client's follow-up init() instead of failing
  in the gap between disconnect() and init().

Refs: TAT-4041
Retry rate-limited pre-order HTTP reads with jittered backoff.
Keep only the controller change: an action waiting on disconnect() waits a
bounded time for the client's follow-up init(), and fails with
PROVIDER_LIFECYCLE_STALE if the account, network or provider changed.

The HTTP order-path reads, the 429 retry and the getMaxLeverage change move
to a stacked follow-up so the transport choice can be validated separately.

Refs: TAT-4041
Pre-order reads used the WebSocket info client, which is unavailable for the
whole reconnect even though the HTTP exchange client can still take the order.

- Read the margin-mode lock, HIP-3 balances and spot metadata, unified-account
  and referral setup, asset-map rebuilds and the updatePositionTPSL open-order
  read over HTTP, and check the HTTP client in the #ensureReady gate.
- Drop the WebSocket gate from getMaxLeverage, whose metadata read is already
  HTTP, so orders are no longer capped at the 3x fallback during a reconnect.
- Retry rate-limited (429) pre-order HTTP reads up to twice with jittered
  exponential backoff.

Refs: TAT-4041
Base automatically changed from TAT-4041-fix-fix-order-submit-ws-reconnect to main September 30, 2026 00:51
pull Bot pushed a commit to Reality2byte/core that referenced this pull request Sep 30, 2026
…t reconnect (MetaMask#10589)

## Explanation

Mobile reconnects Perps with `PerpsController.disconnect()`, a short
cleanup delay, then `init()`. That covers a foreground resume after a
failed ping, and account or network changes.
`#getActiveProviderWhenReady` waited for the disconnect but not for the
`init()` that follows it. An order submitted in that window woke up to
an uninitialized controller and failed with `CLIENT_NOT_INITIALIZED`.

This PR changes only the controller:

- If an action waited on a `disconnect()` and the controller is still
uninitialized afterwards, it waits up to
`PERPS_CONSTANTS.ConnectionTimeoutMs` for the client's `init()` to
start, then runs. If no `init()` follows, it fails as before. The
controller never starts a connection itself.
- If the reconnect switched the selected account, the network or the
active provider, the action fails with `PROVIDER_LIFECYCLE_STALE`
instead of running under the new context. The context is recorded when
the action is issued.

Behavior change: after a final disconnect with no follow-up `init()`, an
action that was already in flight now takes up to 10 s to fail instead
of failing right away. Actions issued after the disconnect are
unaffected.

The provider-side part of TAT-4041 (order-path reads over HTTP during a
WebSocket reconnect, the `getMaxLeverage` 3x cap, 429 retry) is split
into the stacked draft MetaMask#10590. It stays in draft until a transport
stress test settles the long-term transport choice.

Not in this PR:

- `PROVIDER_LIFECYCLE_STALE` for an order already in flight inside the
provider when the account changes. Continuing under the new account
would be wrong. How that code is shown to the user is up to the client.

## Reproduction

The real headless `PerpsController`, SDK transports and signer ran
against HyperLiquid testnet, with a client driver calling the real
`disconnect()` and a delayed `init()`.

| Controlled condition | Unfixed source (`1f15b7c`) | With the
controller change |
| --- | --- | --- |
| Client `disconnect()` followed by delayed `init()` | Disconnect
finished at 2,970 ms, and the order failed with `CLIENT_NOT_INITIALIZED`
at 2,970 ms. `init()` only started at 3,470 ms. | `init()` finished at
5,260 ms. Order `61421723850` succeeded at 7,882 ms, was observed open,
and was canceled by ID. 0 BTC orders and 0 positions after teardown. |

That run was recorded on `9eb852d`, which also contained the provider
changes now in MetaMask#10590. The controller code on this path is the same
apart from narrowing: the wait now only follows a `disconnect()`, not a
reinitialization. I have not rerun the testnet reproduction on
`9149042`. No Mobile or Extension UI was exercised.

## Validation (`9149042`)

- `PerpsController` suites: `8 passed`, `465 passed`.
- The 4 new lifecycle tests all fail on `main`: order placed across
disconnect then init, bounded wait, and refusal after an account switch
or a network switch. With only the context guard removed, just the 2
switch tests fail.
- `oxlint` and `oxfmt` clean; `changelog:validate` passes; the package
build type-checks.

## References

- Fixes
[TAT-4041](https://consensyssoftware.atlassian.net/browse/TAT-4041)
(controller part). Provider part: MetaMask#10590
- Related:
[TAT-4013](https://consensyssoftware.atlassian.net/browse/TAT-4013)
(`PROVIDER_LIFECYCLE_STALE` on account switch),
[TAT-3871](https://consensyssoftware.atlassian.net/browse/TAT-3871)
(retry and idempotency of the exchange call)

## Checklist

- [x] I've updated the test suite for new or updated code as appropriate
- [x] I've updated documentation (JSDoc, Markdown, etc.) for new or
updated code as appropriate
- [x] I've communicated my changes to consumers by [updating changelogs
for packages I've
changed](https://github.com/MetaMask/core/tree/main/docs/processes/updating-changelogs.md)
- [ ] I've introduced [breaking
changes](https://github.com/MetaMask/core/tree/main/docs/processes/breaking-changes.md)
in this PR and have prepared draft pull requests for clients and
consumer packages to resolve them


[TAT-4041]:
https://consensyssoftware.atlassian.net/browse/TAT-4041?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant