Steps to reproduce
This deliberately disconnects a replica tunnel to reproduce the recovery failure.
-
With dstack 0.22.2, create a test gateway with one HTTP service and one replica (auth: false). Confirm its service URL through the gateway returns 200:
SERVICE_URL='http://<service>.<gateway-domain>/'
curl -i --max-time 10 "$SERVICE_URL"
-
On the gateway, find the ssh PID owning the replica.sock listener, then terminate that process:
sudo ss -xlpn | grep '/replica.sock'
sudo kill -KILL <ssh-pid>
-
From the client, wait 60 seconds and request the same service:
sleep 60
curl -i --max-time 10 "$SERVICE_URL"
Expected behaviour
The gateway reconnects to the healthy replica, and requests recover without restarting the gateway.
Actual behaviour
Every probe returned nginx 502 for 60 seconds. The backend still returned 200 directly, and the gateway process stayed running. The Unix socket file remained, but no process was listening on it. Restarting dstack.gateway.service restored HTTP 200.
Notes
ServiceConnectionPool reuses an existing connection without checking whether its background SSH process is alive. Server reconciliation checks replica registrations, not tunnel health. Nginx therefore keeps forwarding to the abandoned Unix socket.
An instrumented load test also reached this state: memory pressure stalled the gateway, and the replica's SSH server disconnected it for missing keepalive replies. Stock load tests did not reproduce this persistent failure; one instead caused an OOM kill and automatic gateway restart. The steps above isolate the recovery defect.
Suggested fix: detect and recreate failed service SSH tunnels in the OSS proxy, updating nginx if the socket path changes.
Steps to reproduce
This deliberately disconnects a replica tunnel to reproduce the recovery failure.
With dstack 0.22.2, create a test gateway with one HTTP service and one replica (
auth: false). Confirm its service URL through the gateway returns 200:On the gateway, find the
sshPID owning thereplica.socklistener, then terminate that process:From the client, wait 60 seconds and request the same service:
sleep 60 curl -i --max-time 10 "$SERVICE_URL"Expected behaviour
The gateway reconnects to the healthy replica, and requests recover without restarting the gateway.
Actual behaviour
Every probe returned nginx 502 for 60 seconds. The backend still returned 200 directly, and the gateway process stayed running. The Unix socket file remained, but no process was listening on it. Restarting
dstack.gateway.servicerestored HTTP 200.Notes
ServiceConnectionPool reuses an existing connection without checking whether its background SSH process is alive. Server reconciliation checks replica registrations, not tunnel health. Nginx therefore keeps forwarding to the abandoned Unix socket.
An instrumented load test also reached this state: memory pressure stalled the gateway, and the replica's SSH server disconnected it for missing keepalive replies. Stock load tests did not reproduce this persistent failure; one instead caused an OOM kill and automatic gateway restart. The steps above isolate the recovery defect.
Suggested fix: detect and recreate failed service SSH tunnels in the OSS proxy, updating nginx if the socket path changes.