Skip to content

Harden hypervisor process liveness checks - #363

Open
yummybomb wants to merge 8 commits into
hypeship/generalize-vgpu-devicefrom
hypeship/hypervisor-liveness
Open

Harden hypervisor process liveness checks#363
yummybomb wants to merge 8 commits into
hypeship/generalize-vgpu-devicefrom
hypeship/hypervisor-liveness

Conversation

@yummybomb

@yummybomb yummybomb commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

Layer 1 of the vendor VFIO vGPU stack (generalize-vgpu-devicethisvendor-vfio-backendvendor-vfio-vgpu). Pure hypervisor-process hardening with no vGPU-specific code; reviewable in isolation.

The upper layers guard vGPU release decisions on "is this instance's hypervisor still alive", so the liveness answer has to be trustworthy first:

  • Unify liveness checks on ProcessExists — one exported, EPERM-aware, zombie-filtering definition instead of scattered bare kill(pid, 0) probes. EPERM means the process exists but cannot be signaled; treating it as dead would be wrong.
  • Wait for non-child hypervisor exit before finishing kill — after a hypeman restart the hypervisor is not our child, so Wait4 returns ECHILD immediately and the kill loop finished before the process had exited. Poll for actual exit in that case.
  • HypervisorProcessExists: verify socket ownership — a bare PID probe treats any process that reused a stored hypervisor PID as the owning VMM. On Linux, require the PID to own the instance's hypervisor socket before reporting it alive.

Testing

  • go build ./..., go vet clean
  • go test -race ./lib/instances/ targeted suites pass (TestCreateInstanceWithNetwork requires image pulls + iptables and fails in this environment on the unmodified base as well)

Note

Medium Risk
Changes delete-time SIGKILL targeting and hypervisor liveness signals; the design favors skipping kills over killing unrelated processes, but mis-resolved socket ownership could still leave a leaked VMM or kill the wrong owner in edge cases.

Overview
Hardens how hypeman decides whether a hypervisor is still alive and which PID to tear down on delete, so stale metadata after a restart cannot point SIGKILL at the wrong process.

Linux ResolveProcessPID now returns a confirmed flag when the PID comes from socket inode / FD ownership (vs cmdline fallback). Socket lookup in /proc/net/unix only treats __SO_ACCEPTCON listeners as the owner, ignores accepted connections that share the path, errors on ambiguous duplicate listeners, and surfaces FD scan failures instead of silently falling through.

Instance layer: unexported processExists becomes exported ProcessExists (EPERM-aware, zombie filtering on Linux). New HypervisorProcessExists requires the stored PID to own the instance control socket when ownership can be confirmed. killHypervisor only kills when ownership is confirmed (or on non-Linux / no socket path); on a confirmed mismatch it kills the socket owner; when ownership is unknown it skips the kill. The post-SIGKILL wait loop keeps polling until the process is gone when Wait4 returns ECHILD (hypervisor not a child after hypeman restart).

Integration tests cover reused PIDs, missing sockets, and socket rebound scenarios.

Reviewed by Cursor Bugbot for commit f8fbe79. Bugbot is set up for automated code reviews on this repo. Configure here.

Comment thread lib/hypervisor/socket_pid_linux.go
@yummybomb
yummybomb force-pushed the hypeship/hypervisor-liveness branch 2 times, most recently from 76b9f42 to 78fc483 Compare August 7, 2026 20:52
kill(pid, 0) returning EPERM means the process exists but cannot be
signaled, and a zombie PID passes a bare kill(0) probe. Export the
EPERM-aware, zombie-filtering processExists helper so every hypervisor
liveness check shares one definition.
After a hypeman restart the hypervisor is not our child, so Wait4
returns ECHILD immediately and the kill loop finished before the
process had exited. Poll for actual process exit in that case.
A bare liveness probe treats any process that reused a stored
hypervisor PID as the owning VMM. Require the PID to own the
instance's hypervisor socket on Linux before reporting it alive.
Accepted server-side sockets appear in /proc/net/unix with the same
bound path as the listener, so any connected API client made
socketRefForPath report multiple inodes and pid-reuse protection fell
back to unconfirmed while the control socket was in use. Only entries
with __SO_ACCEPTCON identify the owning process; duplicate listeners
from unlink-and-rebind still resolve as unconfirmed.
@yummybomb
yummybomb force-pushed the hypeship/hypervisor-liveness branch from 78fc483 to f8fbe79 Compare August 8, 2026 01:05

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit f8fbe79. Configure here.

Comment thread lib/instances/delete.go
log.DebugContext(ctx, "hypervisor process killed", "instance_id", inst.Id, "pid", pid)
break
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Socket unlinked after skipped kill

High Severity

When socket ownership cannot be confirmed or an unconfirmed PID mismatch skips the kill, killHypervisor still calls os.Remove on SocketPath. That unlinks a live VMM’s control socket after the new logic deliberately refused to kill it, turning an intended recoverable leak into an unmanageable orphan.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit f8fbe79. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant