Steps to reproduce
- Run
dstack server from a terminal and create a cloud fleet with one instance.
- Make the instance's SSH master unresponsive or let it die while the server is checking it (e.g., an unstable network link to the instance).
- Submit a run to that fleet.
Actual behaviour
The run stays in submitted indefinitely, and the instance is never processed again: no checks and no idle termination, so it keeps billing.
The server's instance check is stuck in ssh -S <control.sock> -O check ubuntu@<instance>. In our case it hung for 35+ minutes, holding a TCP connection to the instance's port 22 and the server's terminal (/dev/tty), waiting for an interactive login prompt. Killing that ssh process unblocked everything within ~30 seconds.
SSHTunnel.open() runs ssh with -F, -i, BatchMode=yes and timeout=SSH_TIMEOUT, but the control commands used by check(), close(), acheck(), aclose() and aexec() have none of these and run without a timeout:
|
def close_command(self) -> List[str]: |
|
return [self.ssh_exec_path, "-S", self.control_sock_path, "-O", "exit", self.destination] |
|
|
|
def check_command(self) -> List[str]: |
|
return [self.ssh_exec_path, "-S", self.control_sock_path, "-O", "check", self.destination] |
|
|
|
def exec_command(self) -> List[str]: |
|
return [self.ssh_exec_path, "-S", self.control_sock_path, self.destination] |
|
def close(self) -> None: |
|
if not os.path.exists(self.control_sock_path): |
|
logger.debug( |
|
"Control socket does not exist, it seems that ssh process has already exited" |
|
) |
|
return |
|
proc = subprocess.run( |
|
self.close_command(), stdout=subprocess.PIPE, stderr=subprocess.STDOUT |
|
) |
|
if proc.returncode: |
|
logger.error( |
|
"Failed to close SSH tunnel, exit status: %d, output: %s", |
|
proc.returncode, |
|
proc.stdout, |
|
) |
|
|
|
async def aclose(self) -> None: |
|
if not os.path.exists(self.control_sock_path): |
|
logger.debug( |
|
"Control socket does not exist, it seems that ssh process has already exited" |
|
) |
|
return |
|
proc = await asyncio.create_subprocess_exec( |
|
*self.close_command(), stdout=subprocess.PIPE, stderr=subprocess.STDOUT |
|
) |
|
await proc.wait() |
|
if proc.returncode: |
|
logger.error( |
|
"Failed to close SSH tunnel, exit status: %d, output: %s", |
|
proc.returncode, |
|
proc.stdout, |
|
) |
|
|
|
def check(self) -> bool: |
|
proc = subprocess.run( |
|
self.check_command(), stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL |
|
) |
|
return proc.returncode == 0 |
|
|
|
async def acheck(self) -> bool: |
|
proc = await asyncio.create_subprocess_exec( |
|
*self.check_command(), stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL |
|
) |
|
await proc.wait() |
|
ok = proc.returncode == 0 |
|
return ok |
|
|
|
async def aexec(self, command: str) -> str: |
|
proc = await asyncio.create_subprocess_exec( |
|
*self.exec_command(), command, stdout=subprocess.PIPE, stderr=subprocess.PIPE |
|
) |
|
stdout, stderr = await proc.communicate() |
|
if proc.returncode != 0: |
|
raise SSHError(stderr.decode()) |
|
return stdout.decode() |
The ssh client can then hang in two ways:
- If the master accepts the control connection but never replies, the client waits forever: the hello exchange has a timeout only if
ConnectTimeout is set.
https://github.com/openssh/openssh-portable/blob/6849957945754e6551e515f41e8cf3937cda222d/mux.c#L2322-L2331
- If the master breaks off the hello exchange (e.g., it is exiting), the client falls back to a regular connection to the destination, even for
-O commands. Without dstack's identity and BatchMode, that connection uses the user's SSH config and keys and waits for a password or host key confirmation on /dev/tty when the server runs in a terminal.
https://github.com/openssh/openssh-portable/blob/6849957945754e6551e515f41e8cf3937cda222d/ssh.c#L1594-L1600
InstanceConnectionPool.get_or_open() calls check() and close() while holding the per-instance lock, so a hung call blocks every user of that instance's connection. The instance pipeline keeps heartbeating its lock on the instance for the whole time, so the jobs pipeline can't lock the instance for assignment either:
|
key = InstanceConnectionKey.from_jpd(jpd, jrd) |
|
lock = self._get_access_lock(key) |
|
with lock: |
|
if self._closed: |
|
return None |
|
conn = self._connections.get(key) |
|
if conn is not None: |
|
if conn.is_alive(): |
|
return conn |
|
# The master process is gone — evict and reopen. |
|
logger.debug("Instance connection %s is dead, reopening", key) |
|
self._connections.pop(key) |
|
try: |
|
conn.close() |
|
except Exception: |
|
logger.exception("Failed to close instance connection %s", key) |
|
try: |
|
conn = InstanceConnection(ssh_private_key, jpd, jrd) |
|
conn.open() |
|
except SSHError: |
|
# error logged in tunnel |
|
return None |
|
self._connections[key] = conn |
|
return conn |
Expected behaviour
Control commands (-O check, -O exit) and aexec fail within a bounded time and never prompt, so a broken master results in the connection being reopened rather than instance processing being blocked.
dstack version
master (0d578c8), server running from source on macOS, OpenSSH_9.8p1.
Server logs
DEBUG dstack._internal.server.background.pipeline_tasks.jobs_submitted:725
job(3c3a91)qwen38-flash-next-coding-h200-0-0: failed to lock existing fleet instances for assignment
DEBUG dstack._internal.server.background.pipeline_tasks.fleets:416
Failed to lock fleet 3cdf9902-c3f1-4f56-9abf-11914dbb120a instances. The fleet will be processed later.
DEBUG dstack._internal.server.background.pipeline_tasks.base:224
Updating lock_expires_at for items: ['8c41209e-9ebc-4e46-b65b-691943e6eadb']
No error is logged for the stuck check itself.
Additional information
Suggested fix: give close(), check(), acheck(), aclose() and aexec() a timeout (killing the process on expiry) and pass -o BatchMode=yes and -o ConnectTimeout=<n> to the control commands; for aexec(), also redirect stdin (-n). Redirecting stdin alone doesn't prevent the prompt, since ssh reads it from /dev/tty.
Steps to reproduce
dstack serverfrom a terminal and create a cloud fleet with one instance.Actual behaviour
The run stays in
submittedindefinitely, and the instance is never processed again: no checks and no idle termination, so it keeps billing.The server's instance check is stuck in
ssh -S <control.sock> -O check ubuntu@<instance>. In our case it hung for 35+ minutes, holding a TCP connection to the instance's port 22 and the server's terminal (/dev/tty), waiting for an interactive login prompt. Killing thatsshprocess unblocked everything within ~30 seconds.SSHTunnel.open()runssshwith-F,-i,BatchMode=yesandtimeout=SSH_TIMEOUT, but the control commands used bycheck(),close(),acheck(),aclose()andaexec()have none of these and run without a timeout:dstack/src/dstack/_internal/core/services/ssh/tunnel.py
Lines 175 to 182 in 0d578c8
dstack/src/dstack/_internal/core/services/ssh/tunnel.py
Lines 222 to 276 in 0d578c8
The
sshclient can then hang in two ways:ConnectTimeoutis set.https://github.com/openssh/openssh-portable/blob/6849957945754e6551e515f41e8cf3937cda222d/mux.c#L2322-L2331
-Ocommands. Without dstack's identity andBatchMode, that connection uses the user's SSH config and keys and waits for a password or host key confirmation on/dev/ttywhen the server runs in a terminal.https://github.com/openssh/openssh-portable/blob/6849957945754e6551e515f41e8cf3937cda222d/ssh.c#L1594-L1600
InstanceConnectionPool.get_or_open()callscheck()andclose()while holding the per-instance lock, so a hung call blocks every user of that instance's connection. The instance pipeline keeps heartbeating its lock on the instance for the whole time, so the jobs pipeline can't lock the instance for assignment either:dstack/src/dstack/_internal/server/services/runner/pool.py
Lines 95 to 118 in 0d578c8
Expected behaviour
Control commands (
-O check,-O exit) andaexecfail within a bounded time and never prompt, so a broken master results in the connection being reopened rather than instance processing being blocked.dstack version
master (0d578c8), server running from source on macOS, OpenSSH_9.8p1.
Server logs
DEBUG dstack._internal.server.background.pipeline_tasks.jobs_submitted:725 job(3c3a91)qwen38-flash-next-coding-h200-0-0: failed to lock existing fleet instances for assignment DEBUG dstack._internal.server.background.pipeline_tasks.fleets:416 Failed to lock fleet 3cdf9902-c3f1-4f56-9abf-11914dbb120a instances. The fleet will be processed later. DEBUG dstack._internal.server.background.pipeline_tasks.base:224 Updating lock_expires_at for items: ['8c41209e-9ebc-4e46-b65b-691943e6eadb']No error is logged for the stuck check itself.
Additional information
Suggested fix: give
close(),check(),acheck(),aclose()andaexec()a timeout (killing the process on expiry) and pass-o BatchMode=yesand-o ConnectTimeout=<n>to the control commands; foraexec(), also redirect stdin (-n). Redirecting stdin alone doesn't prevent the prompt, sincesshreads it from/dev/tty.