Skip to content

How to run with BNB 4bit or 8bit quantization? #3

Description

@fireicewolf

I tryed to modify your example code to run this model on lowvram card by BNB 4bit or 8bit quantization config.

While use bnb 4bit config like below:

qnt_config = BitsAndBytesConfig(load_in_4bit=True,
                                bnb_4bit_quant_type="nf4",
                                bnb_4bit_compute_dtype=torch.float16,
                                bnb_4bit_use_double_quant=True)

First time this issue occured while pixel_values = pixel_values.to(torch.bfloat16).unsqueeze(0)
RuntimeError: Input type (CUDABFloat16Type) and weight type (torch.cuda.HalfTensor) should be the same
Then I changed it to pixel_values = pixel_values.to(llm_dtype).unsqueeze(0)(llm_dtype is llava models weight load dtype)
RuntimeError: self and mat2 must have the same dtype, but got Half and Byte

these error should be caused by image input dtype.

Any idea to make it works?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions