Skip to content

Abstract vGPU devices behind a framework dispatch - #366

Open
yummybomb wants to merge 10 commits into
mainfrom
hypeship/vgpu-framework-abstraction
Open

Abstract vGPU devices behind a framework dispatch#366
yummybomb wants to merge 10 commits into
mainfrom
hypeship/vgpu-framework-abstraction

Conversation

@yummybomb

@yummybomb yummybomb commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

Bottom layer of the vendor VFIO vGPU stack (this#322#363#364#321). Behavior-preserving refactor only — no lifecycle semantics change in this layer.

Kernel 6.8 hosts assign NVIDIA vGPUs through a vendor-specific VFIO interface instead of mdev, so the mdev-shaped seams get generalized before the new backend lands above:

  • VGPUDevice / VGPUFramework abstraction — mdev moves behind a framework dispatch (CreateVGPU / DestroyVGPU / DiscoverVGPU), with assignments described by a framework + device path instead of a bare mdev UUID.
  • Hypervisor selection — vGPU instances require QEMU; the policy moves to callers instead of being implied by mdev.
  • Metadata — VF allocation tracked with a single field; instance metadata carries GPUFramework / GPUDevicePath alongside the mdev UUID.
  • QEMU args — vGPU attaches via sysfsdev, and the dead mdev branch is dropped from PCI passthrough args.

Testing

  • go build ./..., go vet clean
  • go test -race ./lib/devices/ ./lib/hypervisor/... and targeted lib/instances suites pass (TestSocketCacheKeyChangesWhenSocketIsRecreated and the network/image-dependent instances tests fail identically on the unmodified stack head in this environment)

Note

Medium Risk
Changes GPU assignment lifecycle, persisted metadata, and hypervisor passthrough wiring; delete now blocks on failed vGPU release instead of best-effort cleanup.

Overview
Introduces a vGPU framework layer (VGPUFramework, VGPUDevice, CreateVGPU / DestroyVGPU) so mdev is one backend today and future assignment styles can plug in without rewriting instance code. SR-IOV VF occupancy is exposed as allocated instead of has_mdev.

Instance lifecycle now persists GPUFramework and GPUDevicePath (with legacy GPUMdevUUID fallback for paths), routes create/start/stop/delete through releaseStoredVGPU, and wires the hypervisor via dedicated VGPUDevicePath rather than stuffing mdev sysfs into PCIDevices. Delete fails and keeps metadata if vGPU teardown fails (so a compatible version can retry); stop logs and retains assignment metadata on failure; start clears stale assignments before boot.

QEMU and Cloud Hypervisor always attach vGPUs with vfio-pci,sysfsdev=… from VGPUDevicePath. GPU.md adds rollback steps before downgrading off the active vGPU framework.

Reviewed by Cursor Bugbot for commit f677357. Bugbot is set up for automated code reviews on this repo. Configure here.

@yummybomb yummybomb changed the title hypeship/vgpu framework abstraction Abstract vGPU devices behind a framework dispatch Aug 6, 2026
@yummybomb
yummybomb marked this pull request as ready for review August 6, 2026 19:20
@yummybomb
yummybomb force-pushed the hypeship/vgpu-framework-abstraction branch from 284269f to c3d6a2f Compare August 6, 2026 19:26
@yummybomb
yummybomb force-pushed the hypeship/vgpu-framework-abstraction branch from c3d6a2f to f677357 Compare August 6, 2026 19:40

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit f677357. Configure here.

Comment thread lib/instances/start.go
if err := releaseStoredVGPU(ctx, stored); err != nil {
log.ErrorContext(ctx, "failed to release stale vGPU before start", "instance_id", id, "error", err)
return nil, fmt.Errorf("release stale vGPU before start: %w", err)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fork start can destroy source vGPU

Medium Severity

start now calls releaseStoredVGPU, which destroys whatever assignment is recorded in instance metadata. Stop also retains that assignment when destroy fails. Forks still copy GPUFramework, GPUDevicePath, and GPUMdevUUID from the source (unlike network identity, which is cleared), so starting such a fork can destroy the source’s still-tracked vGPU.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit f677357. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant