Skip to content

Backend fails to start when torch cannot be loaded #730

Description

@thcp

Split out of #723.

What goes wrong

When import torch raises anything other than ImportError, the backend does not start. The reporter in #723 hit OSError: [WinError 127] ... Error loading c10_cuda.dll, and startup died with exit code 1 before the app could open.

With the device setting forced to CPU the backend does start, but every GET /api/settings then returns 500, so the Settings page cannot load its device list.

Why

available_torch_devices() in app/core/config.py only catches ImportError. A torch whose DLLs cannot load raises OSError. app/main.py also probes the device at import time to log it, so an "auto" device setting turns that exception into a startup crash.

Constraints

  • CPU always works, so the app should start on CPU and say the GPU runtime is damaged, not die.
  • The failure must stay visible. Falling back silently would leave a user wondering why separation is slow, or why every job fails.

Activity

  1. added
    bugSomething isn't working
    on Sep 30, 2026
  2. self-assigned this
    on Oct 1, 2026
  3. added a commit that references this issue on Oct 1, 2026
    d78fa34
  4. added a commit that references this issue on Oct 2, 2026
    939d66e
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions