I tryed to modify your example code to run this model on lowvram card by BNB 4bit or 8bit quantization config.
While use bnb 4bit config like below:
qnt_config = BitsAndBytesConfig(load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16,
bnb_4bit_use_double_quant=True)
First time this issue occured while pixel_values = pixel_values.to(torch.bfloat16).unsqueeze(0)
RuntimeError: Input type (CUDABFloat16Type) and weight type (torch.cuda.HalfTensor) should be the same
Then I changed it to pixel_values = pixel_values.to(llm_dtype).unsqueeze(0)(llm_dtype is llava models weight load dtype)
RuntimeError: self and mat2 must have the same dtype, but got Half and Byte
these error should be caused by image input dtype.
Any idea to make it works?
I tryed to modify your example code to run this model on lowvram card by BNB 4bit or 8bit quantization config.
While use bnb 4bit config like below:
First time this issue occured while
pixel_values = pixel_values.to(torch.bfloat16).unsqueeze(0)RuntimeError: Input type (CUDABFloat16Type) and weight type (torch.cuda.HalfTensor) should be the sameThen I changed it to
pixel_values = pixel_values.to(llm_dtype).unsqueeze(0)(llm_dtype is llava models weight load dtype)RuntimeError: self and mat2 must have the same dtype, but got Half and Bytethese error should be caused by image input dtype.
Any idea to make it works?