Skip to content

black-forest-labs/FLUX.2-klein-9B not working with lora with lokr #13261

Description

@chaowenguo

Describe the bug

Traceback (most recent call last):
File "/usr/local/lib/python3.11/site-packages/aiohttp/web_protocol.py", line 510, in _handle_request
resp = await request_handler(request)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/aiohttp/web_app.py", line 569, in _handle
return await handler(request)
^^^^^^^^^^^^^^^^^^^^^^
File "/home/studio_service/PROJECT/app.py", line 30, in imageGet
request.app.pop('future').result().save(buffer, 'png')
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/concurrent/futures/thread.py", line 58, in run
result = self.fn(*self.args, **self.kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/studio_service/PROJECT/app.py", line 16, in imageGenerate
flux.load_lora_weights(modelscope.model_file_download('chaowenguo/lora', 'klein_snofs_v1_1.safetensors'), adapter_name='snofs')
File "/usr/local/lib/python3.11/site-packages/diffusers/loaders/lora_pipeline.py", line 5713, in load_lora_weights
state_dict, metadata = self.lora_state_dict(pretrained_model_name_or_path_or_dict, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/huggingface_hub/utils/_validators.py", line 114, in _inner_fn
return fn(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/diffusers/loaders/lora_pipeline.py", line 5682, in lora_state_dict
state_dict = _convert_non_diffusers_flux2_lora_to_diffusers(state_dict)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/diffusers/loaders/lora_conversion_utils.py", line 2435, in _convert_non_diffusers_flux2_lora_to_diffusers
raise ValueError(f"original_state_dict should be empty at this point but has {original_state_dict.keys()=}.")
ValueError: original_state_dict should be empty at this point but has original_state_dict.keys()=dict_keys(['double_blocks.0.img_attn.proj.alpha', 'double_blocks.0.img_attn.proj.lokr_w1', 'double_blocks.0.img_attn.proj.lokr_w2', 'double_blocks.0.img_attn.qkv.alpha', 'double_blocks.0.img_attn.qkv.lokr_w1', 'double_blocks.0.img_attn.qkv.lokr_w2', 'double_blocks.0.img_mlp.0.alpha', 'double_blocks.0.img_mlp.0.lokr_w1', 'double_blocks.0.img_mlp.0.lokr_w2', 'double_blocks.0.img_mlp.2.alpha', 'double_blocks.0.img_mlp.2.lokr_w1', 'double_blocks.0.img_mlp.2.lokr_w2', 'double_blocks.0.txt_attn.proj.alpha', 'double_blocks.0.txt_attn.proj.lokr_w1', 'double_blocks.0.txt_attn.proj.lokr_w2', 'double_blocks.0.txt_attn.qkv.alpha', 'double_blocks.0.txt_attn.qkv.lokr_w1', 'double_blocks.0.txt_attn.qkv.lokr_w2', 'double_blocks.0.txt_mlp.0.alpha', 'double_blocks.0.txt_mlp.0.lokr_w1', 'double_blocks.0.txt_mlp.0.lokr_w2', 'double_blocks.0.txt_mlp.2.alpha', 'double_blocks.0.txt_mlp.2.lokr_w1', 'double_blocks.0.txt_mlp.2.lokr_w2', 'double_blocks.1.img_attn.proj.alpha', 'double_blocks.1.img_attn.proj.lokr_w1', 'double_blocks.1.img_attn.proj.lokr_w2', 'double_blocks.1.img_attn.qkv.alpha', 'double_blocks.1.img_attn.qkv.lokr_w1', 'double_blocks.1.img_attn.qkv.lokr_w2', 'double_blocks.1.img_mlp.0.alpha', 'double_blocks.1.img_mlp.0.lokr_w1', 'double_blocks.1.img_mlp.0.lokr_w2', 'double_blocks.1.img_mlp.2.alpha', 'double_blocks.1.img_mlp.2.lokr_w1', 'double_blocks.1.img_mlp.2.lokr_w2', 'double_blocks.1.txt_attn.proj.alpha', 'double_blocks.1.txt_attn.proj.lokr_w1', 'double_blocks.1.txt_attn.proj.lokr_w2', 'double_blocks.1.txt_attn.qkv.alpha', 'double_blocks.1.txt_attn.qkv.lokr_w1', 'double_blocks.1.txt_attn.qkv.lokr_w2', 'double_blocks.1.txt_mlp.0.alpha', 'double_blocks.1.txt_mlp.0.lokr_w1', 'double_blocks.1.txt_mlp.0.lokr_w2', 'double_blocks.1.txt_mlp.2.alpha', 'double_blocks.1.txt_mlp.2.lokr_w1', 'double_blocks.1.txt_mlp.2.lokr_w2', 'double_blocks.2.img_attn.proj.alpha', 'double_blocks.2.img_attn.proj.lokr_w1', 'double_blocks.2.img_attn.proj.lokr_w2', 'double_blocks.2.img_attn.qkv.alpha', 'double_blocks.2.img_attn.qkv.lokr_w1', 'double_blocks.2.img_attn.qkv.lokr_w2', 'double_blocks.2.img_mlp.0.alpha', 'double_blocks.2.img_mlp.0.lokr_w1', 'double_blocks.2.img_mlp.0.lokr_w2', 'double_blocks.2.img_mlp.2.alpha', 'double_blocks.2.img_mlp.2.lokr_w1', 'double_blocks.2.img_mlp.2.lokr_w2', 'double_blocks.2.txt_attn.proj.alpha', 'double_blocks.2.txt_attn.proj.lokr_w1', 'double_blocks.2.txt_attn.proj.lokr_w2', 'double_blocks.2.txt_attn.qkv.alpha', 'double_blocks.2.txt_attn.qkv.lokr_w1', 'double_blocks.2.txt_attn.qkv.lokr_w2', 'double_blocks.2.txt_mlp.0.alpha', 'double_blocks.2.txt_mlp.0.lokr_w1', 'double_blocks.2.txt_mlp.0.lokr_w2', 'double_blocks.2.txt_mlp.2.alpha', 'double_blocks.2.txt_mlp.2.lokr_w1', 'double_blocks.2.txt_mlp.2.lokr_w2', 'double_blocks.3.img_attn.proj.alpha', 'double_blocks.3.img_attn.proj.lokr_w1', 'double_blocks.3.img_attn.proj.lokr_w2', 'double_blocks.3.img_attn.qkv.alpha', 'double_blocks.3.img_attn.qkv.lokr_w1', 'double_blocks.3.img_attn.qkv.lokr_w2', 'double_blocks.3.img_mlp.0.alpha', 'double_blocks.3.img_mlp.0.lokr_w1', 'double_blocks.3.img_mlp.0.lokr_w2', 'double_blocks.3.img_mlp.2.alpha', 'double_blocks.3.img_mlp.2.lokr_w1', 'double_blocks.3.img_mlp.2.lokr_w2', 'double_blocks.3.txt_attn.proj.alpha', 'double_blocks.3.txt_attn.proj.lokr_w1', 'double_blocks.3.txt_attn.proj.lokr_w2', 'double_blocks.3.txt_attn.qkv.alpha', 'double_blocks.3.txt_attn.qkv.lokr_w1', 'double_blocks.3.txt_attn.qkv.lokr_w2', 'double_blocks.3.txt_mlp.0.alpha', 'double_blocks.3.txt_mlp.0.lokr_w1', 'double_blocks.3.txt_mlp.0.lokr_w2', 'double_blocks.3.txt_mlp.2.alpha', 'double_blocks.3.txt_mlp.2.lokr_w1', 'double_blocks.3.txt_mlp.2.lokr_w2', 'double_blocks.4.img_attn.proj.alpha', 'double_blocks.4.img_attn.proj.lokr_w1', 'double_blocks.4.img_attn.proj.lokr_w2', 'double_blocks.4.img_attn.qkv.alpha', 'double_blocks.4.img_attn.qkv.lokr_w1', 'double_blocks.4.img_attn.qkv.lokr_w2', 'double_blocks.4.img_mlp.0.alpha', 'double_blocks.4.img_mlp.0.lokr_w1', 'double_blocks.4.img_mlp.0.lokr_w2', 'double_blocks.4.img_mlp.2.alpha', 'double_blocks.4.img_mlp.2.lokr_w1', 'double_blocks.4.img_mlp.2.lokr_w2', 'double_blocks.4.txt_attn.proj.alpha', 'double_blocks.4.txt_attn.proj.lokr_w1', 'double_blocks.4.txt_attn.proj.lokr_w2', 'double_blocks.4.txt_attn.qkv.alpha', 'double_blocks.4.txt_attn.qkv.lokr_w1', 'double_blocks.4.txt_attn.qkv.lokr_w2', 'double_blocks.4.txt_mlp.0.alpha', 'double_blocks.4.txt_mlp.0.lokr_w1', 'double_blocks.4.txt_mlp.0.lokr_w2', 'double_blocks.4.txt_mlp.2.alpha', 'double_blocks.4.txt_mlp.2.lokr_w1', 'double_blocks.4.txt_mlp.2.lokr_w2', 'double_blocks.5.img_attn.proj.alpha', 'double_blocks.5.img_attn.proj.lokr_w1', 'double_blocks.5.img_attn.proj.lokr_w2', 'double_blocks.5.img_attn.qkv.alpha', 'double_blocks.5.img_attn.qkv.lokr_w1', 'double_blocks.5.img_attn.qkv.lokr_w2', 'double_blocks.5.img_mlp.0.alpha', 'double_blocks.5.img_mlp.0.lokr_w1', 'double_blocks.5.img_mlp.0.lokr_w2', 'double_blocks.5.img_mlp.2.alpha', 'double_blocks.5.img_mlp.2.lokr_w1', 'double_blocks.5.img_mlp.2.lokr_w2', 'double_blocks.5.txt_attn.proj.alpha', 'double_blocks.5.txt_attn.proj.lokr_w1', 'double_blocks.5.txt_attn.proj.lokr_w2', 'double_blocks.5.txt_attn.qkv.alpha', 'double_blocks.5.txt_attn.qkv.lokr_w1', 'double_blocks.5.txt_attn.qkv.lokr_w2', 'double_blocks.5.txt_mlp.0.alpha', 'double_blocks.5.txt_mlp.0.lokr_w1', 'double_blocks.5.txt_mlp.0.lokr_w2', 'double_blocks.5.txt_mlp.2.alpha', 'double_blocks.5.txt_mlp.2.lokr_w1', 'double_blocks.5.txt_mlp.2.lokr_w2', 'double_blocks.6.img_attn.proj.alpha', 'double_blocks.6.img_attn.proj.lokr_w1', 'double_blocks.6.img_attn.proj.lokr_w2', 'double_blocks.6.img_attn.qkv.alpha', 'double_blocks.6.img_attn.qkv.lokr_w1', 'double_blocks.6.img_attn.qkv.lokr_w2', 'double_blocks.6.img_mlp.0.alpha', 'double_blocks.6.img_mlp.0.lokr_w1', 'double_blocks.6.img_mlp.0.lokr_w2', 'double_blocks.6.img_mlp.2.alpha', 'double_blocks.6.img_mlp.2.lokr_w1', 'double_blocks.6.img_mlp.2.lokr_w2', 'double_blocks.6.txt_attn.proj.alpha', 'double_blocks.6.txt_attn.proj.lokr_w1', 'double_blocks.6.txt_attn.proj.lokr_w2', 'double_blocks.6.txt_attn.qkv.alpha', 'double_blocks.6.txt_attn.qkv.lokr_w1', 'double_blocks.6.txt_attn.qkv.lokr_w2', 'double_blocks.6.txt_mlp.0.alpha', 'double_blocks.6.txt_mlp.0.lokr_w1', 'double_blocks.6.txt_mlp.0.lokr_w2', 'double_blocks.6.txt_mlp.2.alpha', 'double_blocks.6.txt_mlp.2.lokr_w1', 'double_blocks.6.txt_mlp.2.lokr_w2', 'double_blocks.7.img_attn.proj.alpha', 'double_blocks.7.img_attn.proj.lokr_w1', 'double_blocks.7.img_attn.proj.lokr_w2', 'double_blocks.7.img_attn.qkv.alpha', 'double_blocks.7.img_attn.qkv.lokr_w1', 'double_blocks.7.img_attn.qkv.lokr_w2', 'double_blocks.7.img_mlp.0.alpha', 'double_blocks.7.img_mlp.0.lokr_w1', 'double_blocks.7.img_mlp.0.lokr_w2', 'double_blocks.7.img_mlp.2.alpha', 'double_blocks.7.img_mlp.2.lokr_w1', 'double_blocks.7.img_mlp.2.lokr_w2', 'double_blocks.7.txt_attn.proj.alpha', 'double_blocks.7.txt_attn.proj.lokr_w1', 'double_blocks.7.txt_attn.proj.lokr_w2', 'double_blocks.7.txt_attn.qkv.alpha', 'double_blocks.7.txt_attn.qkv.lokr_w1', 'double_blocks.7.txt_attn.qkv.lokr_w2', 'double_blocks.7.txt_mlp.0.alpha', 'double_blocks.7.txt_mlp.0.lokr_w1', 'double_blocks.7.txt_mlp.0.lokr_w2', 'double_blocks.7.txt_mlp.2.alpha', 'double_blocks.7.txt_mlp.2.lokr_w1', 'double_blocks.7.txt_mlp.2.lokr_w2', 'single_blocks.0.linear1.alpha', 'single_blocks.0.linear1.lokr_w1', 'single_blocks.0.linear1.lokr_w2', 'single_blocks.0.linear2.alpha', 'single_blocks.0.linear2.lokr_w1', 'single_blocks.0.linear2.lokr_w2', 'single_blocks.1.linear1.alpha', 'single_blocks.1.linear1.lokr_w1', 'single_blocks.1.linear1.lokr_w2', 'single_blocks.1.linear2.alpha', 'single_blocks.1.linear2.lokr_w1', 'single_blocks.1.linear2.lokr_w2', 'single_blocks.10.linear1.alpha', 'single_blocks.10.linear1.lokr_w1', 'single_blocks.10.linear1.lokr_w2', 'single_blocks.10.linear2.alpha', 'single_blocks.10.linear2.lokr_w1', 'single_blocks.10.linear2.lokr_w2', 'single_blocks.11.linear1.alpha', 'single_blocks.11.linear1.lokr_w1', 'single_blocks.11.linear1.lokr_w2', 'single_blocks.11.linear2.alpha', 'single_blocks.11.linear2.lokr_w1', 'single_blocks.11.linear2.lokr_w2', 'single_blocks.12.linear1.alpha', 'single_blocks.12.linear1.lokr_w1', 'single_blocks.12.linear1.lokr_w2', 'single_blocks.12.linear2.alpha', 'single_blocks.12.linear2.lokr_w1', 'single_blocks.12.linear2.lokr_w2', 'single_blocks.13.linear1.alpha', 'single_blocks.13.linear1.lokr_w1', 'single_blocks.13.linear1.lokr_w2', 'single_blocks.13.linear2.alpha', 'single_blocks.13.linear2.lokr_w1', 'single_blocks.13.linear2.lokr_w2', 'single_blocks.14.linear1.alpha', 'single_blocks.14.linear1.lokr_w1', 'single_blocks.14.linear1.lokr_w2', 'single_blocks.14.linear2.alpha', 'single_blocks.14.linear2.lokr_w1', 'single_blocks.14.linear2.lokr_w2', 'single_blocks.15.linear1.alpha', 'single_blocks.15.linear1.lokr_w1', 'single_blocks.15.linear1.lokr_w2', 'single_blocks.15.linear2.alpha', 'single_blocks.15.linear2.lokr_w1', 'single_blocks.15.linear2.lokr_w2', 'single_blocks.16.linear1.alpha', 'single_blocks.16.linear1.lokr_w1', 'single_blocks.16.linear1.lokr_w2', 'single_blocks.16.linear2.alpha', 'single_blocks.16.linear2.lokr_w1', 'single_blocks.16.linear2.lokr_w2', 'single_blocks.17.linear1.alpha', 'single_blocks.17.linear1.lokr_w1', 'single_blocks.17.linear1.lokr_w2', 'single_blocks.17.linear2.alpha', 'single_blocks.17.linear2.lokr_w1', 'single_blocks.17.linear2.lokr_w2', 'single_blocks.18.linear1.alpha', 'single_blocks.18.linear1.lokr_w1', 'single_blocks.18.linear1.lokr_w2', 'single_blocks.18.linear2.alpha', 'single_blocks.18.linear2.lokr_w1', 'single_blocks.18.linear2.lokr_w2', 'single_blocks.19.linear1.alpha', 'single_blocks.19.linear1.lokr_w1', 'single_blocks.19.linear1.lokr_w2', 'single_blocks.19.linear2.alpha', 'single_blocks.19.linear2.lokr_w1', 'single_blocks.19.linear2.lokr_w2', 'single_blocks.2.linear1.alpha', 'single_blocks.2.linear1.lokr_w1', 'single_blocks.2.linear1.lokr_w2', 'single_blocks.2.linear2.alpha', 'single_blocks.2.linear2.lokr_w1', 'single_blocks.2.linear2.lokr_w2', 'single_blocks.20.linear1.alpha', 'single_blocks.20.linear1.lokr_w1', 'single_blocks.20.linear1.lokr_w2', 'single_blocks.20.linear2.alpha', 'single_blocks.20.linear2.lokr_w1', 'single_blocks.20.linear2.lokr_w2', 'single_blocks.21.linear1.alpha', 'single_blocks.21.linear1.lokr_w1', 'single_blocks.21.linear1.lokr_w2', 'single_blocks.21.linear2.alpha', 'single_blocks.21.linear2.lokr_w1', 'single_blocks.21.linear2.lokr_w2', 'single_blocks.22.linear1.alpha', 'single_blocks.22.linear1.lokr_w1', 'single_blocks.22.linear1.lokr_w2', 'single_blocks.22.linear2.alpha', 'single_blocks.22.linear2.lokr_w1', 'single_blocks.22.linear2.lokr_w2', 'single_blocks.23.linear1.alpha', 'single_blocks.23.linear1.lokr_w1', 'single_blocks.23.linear1.lokr_w2', 'single_blocks.23.linear2.alpha', 'single_blocks.23.linear2.lokr_w1', 'single_blocks.23.linear2.lokr_w2', 'single_blocks.3.linear1.alpha', 'single_blocks.3.linear1.lokr_w1', 'single_blocks.3.linear1.lokr_w2', 'single_blocks.3.linear2.alpha', 'single_blocks.3.linear2.lokr_w1', 'single_blocks.3.linear2.lokr_w2', 'single_blocks.4.linear1.alpha', 'single_blocks.4.linear1.lokr_w1', 'single_blocks.4.linear1.lokr_w2', 'single_blocks.4.linear2.alpha', 'single_blocks.4.linear2.lokr_w1', 'single_blocks.4.linear2.lokr_w2', 'single_blocks.5.linear1.alpha', 'single_blocks.5.linear1.lokr_w1', 'single_blocks.5.linear1.lokr_w2', 'single_blocks.5.linear2.alpha', 'single_blocks.5.linear2.lokr_w1', 'single_blocks.5.linear2.lokr_w2', 'single_blocks.6.linear1.alpha', 'single_blocks.6.linear1.lokr_w1', 'single_blocks.6.linear1.lokr_w2', 'single_blocks.6.linear2.alpha', 'single_blocks.6.linear2.lokr_w1', 'single_blocks.6.linear2.lokr_w2', 'single_blocks.7.linear1.alpha', 'single_blocks.7.linear1.lokr_w1', 'single_blocks.7.linear1.lokr_w2', 'single_blocks.7.linear2.alpha', 'single_blocks.7.linear2.lokr_w1', 'single_blocks.7.linear2.lokr_w2', 'single_blocks.8.linear1.alpha', 'single_blocks.8.linear1.lokr_w1', 'single_blocks.8.linear1.lokr_w2', 'single_blocks.8.linear2.alpha', 'single_blocks.8.linear2.lokr_w1', 'single_blocks.8.linear2.lokr_w2', 'single_blocks.9.linear1.alpha', 'single_blocks.9.linear1.lokr_w1', 'single_blocks.9.linear1.lokr_w2', 'single_blocks.9.linear2.alpha', 'single_blocks.9.linear2.lokr_w1', 'single_blocks.9.linear2.lokr_w2']).
[2026-3-12 12:01:21] FC Invoke Start RequestId: 988e5f0a-c97a-4f35-b8b3-331b48835cfb
[2026-3-12 12:01:21] FC Invoke End RequestId: 988e5f0a-c97a-4f35-b8b3-331b48835cfb

Reproduction

import diffusers, torch

diffusers.Flux2KleinPipeline.from_pretrained('black-forest-labs/FLUX.2-klein-9B', torch_dtype=torch.bfloat16, quantization_config=diffusers.PipelineQuantizationConfig(quant_backend='bitsandbytes_4bit', quant_kwargs={'load_in_4bit':True, 'bnb_4bit_quant_type':'nf4', 'bnb_4bit_compute_dtype':torch.bfloat16}, components_to_quantize=['transformer', 'text_encoder'])).save_pretrained('flux')

flux = diffusers.Flux2KleinPipeline.from_pretrained('flux', torch_dtype=torch.bfloat16)
flux.load_lora_weights('chaowenguo/lora', weight_name='klein_snofs_v1_1.safetensors', adapter_name='snofs') #you can find the klein_snofs_v1_1.safetensors in https://www.modelscope.cn/models/chaowenguo/lora/files
flux.set_adapters('snofs', adapter_weights=1)
flux._exclude_from_cpu_offload = ['vae']
flux.enable_model_cpu_offload()
flux(prompt='a gorgeous japanese girl', height=1280, width=720, num_inference_steps=8, guidance_scale=1).images[0]

flux.load_lora_weights('chaowenguo/lora', weight_name='klein_snofs_v1_1.safetensors', adapter_name='snofs') this cause the problem

Logs

System Info

diffusers 0.37.0
ubuntu22.04-cuda12.4.0-py311-torch2.8.0

Who can help?

@sayakpaul

Activity

sayakpaul commented on Mar 12, 2026

@sayakpaul
Member

Did it work with any earlier version of diffusers?

chaowenguo commented on Mar 12, 2026

@chaowenguo
Author

@sayakpaul no v.0.37.0 is the lowest version that support black-forest-labs/FLUX.2-klein-9B

sayakpaul commented on Mar 12, 2026

@sayakpaul
Member

Oh I meant any particular commit where you found it to be working.

chaowenguo commented on Mar 12, 2026

@chaowenguo
Author

sayakpaul commented on Mar 12, 2026

@sayakpaul
Member

I will take a look thanks for the updates.

self-assigned this
on Mar 12, 2026

iwr-redmond commented on Mar 12, 2026

@iwr-redmond

@sayakpaul #13250 may provide the start of a fix for this issue. It resolves a similar problem with Z-Image conversion.

chaowenguo commented on Mar 12, 2026

@chaowenguo
Author

iwr-redmond commented on Mar 12, 2026

@iwr-redmond

It won't fix this exact issue, as the PR is for Z-Image LoKr conversion.

sayakpaul commented on Mar 13, 2026

@sayakpaul
Member

I tried with the following code:

import diffusers, torch

flux = diffusers.Flux2KleinPipeline.from_pretrained(
    'black-forest-labs/FLUX.2-klein-9B', torch_dtype=torch.bfloat16
).to("cuda")
flux.load_lora_weights(
    'chaowenguo/lora', weight_name='klein_snofs_v1_1.safetensors', adapter_name='snofs'
)

And got:

OSError: chaowenguo/lora is not a local folder and is not a valid model identifier listed on 'https://huggingface.co/models'
If this is a private repository, make sure to pass a token having permission to this repo with `token` or log in with `hf auth login`.

Could you please upload the weights on th Hub?

chaowenguo commented on Mar 13, 2026

@chaowenguo
Author

chaowenguo commented on Mar 13, 2026

@chaowenguo
Author

sayakpaul commented on Mar 13, 2026

@sayakpaul
Member

You will have to help us help you. Revisit the issue when you have cleared some space and can provide a snippet that can run without us having to manually downloa things from elsewhere.

chaowenguo commented on Mar 13, 2026

@chaowenguo
Author
import diffusers, torch

diffusers.Flux2KleinPipeline.from_pretrained('black-forest-labs/FLUX.2-klein-9B', torch_dtype=torch.bfloat16, quantization_config=diffusers.PipelineQuantizationConfig(quant_backend='bitsandbytes_4bit', quant_kwargs={'load_in_4bit':True, 'bnb_4bit_quant_type':'nf4', 'bnb_4bit_compute_dtype':torch.bfloat16}, components_to_quantize=['transformer', 'text_encoder'])).save_pretrained('flux')

flux = diffusers.Flux2KleinPipeline.from_pretrained('flux', torch_dtype=torch.bfloat16)
flux.load_lora_weights('puttmorbidly233/lora', weight_name='klein_snofs_v1_2.safetensors', adapter_name='snofs')
flux.set_adapters('snofs', adapter_weights=1)
flux._exclude_from_cpu_offload = ['vae']
flux.enable_model_cpu_offload()
flux(prompt='a gorgeous japanese girl', height=1280, width=720, num_inference_steps=8, guidance_scale=1).images[0]

@sayakpaul

8 remaining items

CalamitousFelicitousness commented on Mar 23, 2026

@CalamitousFelicitousness
Contributor

In this PR I have added support for multiple types of F2 Klein LoRA to sdnext, could be of use. It includes support for LoKR.
Not sure what is the status of this issue @sayakpaul , if nobody is actively working on implementation ping me, I can make a PR.

sayakpaul commented on Mar 23, 2026

@sayakpaul
Member

I saw the merging script disappear, thought it was pretty helpful.

In this PR I have added support for multiple types of F2 Klein LoRA to sdnext, could be of use. It includes support for LoKR.
Not sure what is the status of this issue @sayakpaul , if nobody is actively working on implementation ping me, I can make a PR.

That's generous of you. I was planning to take a look tomorrow. My plan was to support LoKR through peft. Convert the checkpoint to a peft compatible checkpoint and rely on peft's utilities here:

inject_adapter_in_model(
lora_config, self, adapter_name=adapter_name, state_dict=state_dict, **peft_kwargs
)
incompatible_keys = set_peft_model_state_dict(self, state_dict, adapter_name, **peft_kwargs)

just like we do for LoRA.

I am happy to help with your PR if this is the direction you had in mind. But I am open to suggestions.

CalamitousFelicitousness commented on Mar 23, 2026

@CalamitousFelicitousness
Contributor

@sayakpaul
This might get a bit tangential, but I think it might be important to discuss right now.

That's generous of you. I was planning to take a look tomorrow. My plan was to support LoKR through peft. Convert the checkpoint to a peft compatible checkpoint and rely on peft's utilities here:
just like we do for LoRA.
I am happy to help with your PR if this is the direction you had in mind. But I am open to suggestions.

Well, it kind of depends on what's your end goal. One of the main benefits of my solution is that it doesn't care about quantization or what you're stacking on top, LoRA, DoRA, what have you, "it just works".

However, there is also the question of direction in general, having 50 different ways of handling 50 different adaptors is not great.

SDNext uses its own SDNQ quantization engine, so for us quantization support is important. PEFT will not work with that at the moment, so it doesn't have a huge amount of value for us.
From your perspective trying to align to PEFT is natural, since that's your method. But peft's quantization support is LoRA-only for now and backend-specific (the more important bit), not a general solution, unlike delta-based one.

To me it seems like the most prudent way would be to have vector-based method as default with optional path for fuse if it matches one's application. It's more complex though since you have two paths, so I'll await the consensus.

sayakpaul commented on Mar 24, 2026

@sayakpaul
Member

Huh? 🤔

However, there is also the question of direction in general, having 50 different ways of handling 50 different adaptors is not great.

I am not sure if my direction would head that way but I am happy to be shown wrong. It would just be about detecting if the underlying checkpoint has LoKR stuff and the rest of the codepath should largely remain the same. The detection part -- I agree that it is a bit fragile but we also have to keep in mind that we're operating with checkpoints we don't have any control over.

SDNext uses its own SDNQ quantization engine, so for us quantization support is important. PEFT will not work with that at the moment, so it doesn't have a huge amount of value for us.

We test for that stuff at least in a limited capacity:

def test_lora_loading(self):

I think you probably meant LoKR isn't supported with quantization in PEFT. If that's the case, we can work on adding it (cc: @BenjaminBossan).

Having it supported through PEFT is more general and sustainable long-term because:

  • PEFT is widely used and it is also robustly maintained.
  • Training support becomes naturally possible with quantization. This is even more exciting.
  • Properly tested and the integration will also benefit from PEFT's other features.

To me it seems like the most prudent way would be to have vector-based method as default with optional path for fuse if it matches one's application. It's more complex though since you have two paths, so I'll await the consensus.

Fusion is also possible with the methods (fuse_lora(), for example) we expose from the library. If it's not supported, we can work towards supporting it, given its impact as described above.

Happy to work towards these things. But LMK in case I misunderstood something.

CalamitousFelicitousness commented on Mar 24, 2026

@CalamitousFelicitousness
Contributor

Huh? 🤔

However, there is also the question of direction in general, having 50 different ways of handling 50 different adaptors is not great.

I am not sure if my direction would head that way but I am happy to be shown wrong. It would just be about detecting if the underlying checkpoint has LoKR stuff and the rest of the codepath should largely remain the same.

That's what I mean, I was talking about impact of trying to add my method, trying to slap it onto Diffusers rather than using PEFT would mean LoKR path divergence, yours wouldn't.

You can have a look at my crack at PEFT in the commit above, I've only done programatic tests so far.

sayakpaul commented on Mar 24, 2026

@sayakpaul
Member

Makes sense! I guess we can first start with

It would just be about detecting if the underlying checkpoint has LoKR stuff and the rest of the codepath should largely remain the same. The detection part -- I agree that it is a bit fragile but we also have to keep in mind that we're operating with checkpoints we don't have any control over.

Parallely, on the PEFT side, we can start adding support for LoKR x quantization.

I guess this should cover all bases?

CalamitousFelicitousness commented on Mar 24, 2026

@CalamitousFelicitousness
Contributor

I had a crack at it here, didn't have the time to properly test yet, but you could take a glimpse if any of it looks workable. Programatic tests passed.

sayakpaul commented on Mar 24, 2026

@sayakpaul
Member

That looks more than reasonable to me. Please do open it as a PR if you can. I am sure we can iterate on it quite fast. I have some questions about the remapping module which I was hoping to ask on the PR.

BenjaminBossan commented on Mar 24, 2026

@BenjaminBossan
Member

Thanks for the ping. I'm not super familiar with the diffusion ecosystem, so I might be missing some nuances. With that said, this type of problem seems to have 3 components:

  1. Different checkpoint formats for adapters like LoRA, LoKr, etc.
  2. Diffusers only supports LoRA, via PEFT.
  3. PEFT has spotty support for quantization methods.

When it comes to 1, I can't comment on that, it seems to be an eternal issue with no cure.

When it comes to 2, we are aware that this is not ideal. We could try to make Diffusers work with all PEFT methods, but that would be a ton of work and most PEFT methods would probably not find any application, so most of that would be wasted. The approach that we thought of is: Can we convert non-LoRA PEFT adatpers into LoRA PEFT adapters in a generic fashion? IIUC the commit by @CalamitousFelicitousness, it implements a specific LoKr-to-LoRA conversion. In PEFT, we aim at a generic LoRA conversion which works with most PEFT methods, although at the cost that it is a lossy conversion. Initial testing indicates that it works well enough, but more testing and research are needed for sure. (For context, this feature is not only for use with Diffusers, but also other frameworks that only support LoRA, like vLLM).

On top of that, we want to add a benchmark to PEFT to check if methods other than LoRA work well for adapting image generation models. The PR is almost ready (huggingface/peft#3082) but no final results yet. If we find that other methods work well, together with the conversion script, it would unlock the potential to have a wider variety of PEFT methods finding adoption in the community. It goes without saying that having 1:1 conversion scripts for each potential PEFT method is hardly scalable, so generic albeit lossy conversion sounds like the better bet.

Regarding 3., we're mostly adding new quantization methods on demand, so please let us know what is needed. If, say, LoKr + bitsandbytes is required, we can work on it, but we need to know there is demand first. Same if there is a new quantization method that is promising, someone needs to let us know. Of course, with the LoRA conversion approach outlined above, we can take advantage of all the quantization support that already exists for LoRA without having to implement anything specific for another PEFT method.

CalamitousFelicitousness commented on Mar 24, 2026

@CalamitousFelicitousness
Contributor

I like the headway we're making, good comms.

It appears we might have a bit of an interlock with the benchmark; it would be most desirable to make sure the losses suffered due to the conversion are not substantial enough to be of concern. Ultimately if there are no such concerns discreet conversions seem like the wrong way to go.
One thing I'd like to remain mindful of is that even when loss is acceptable in one scenario it might affect another scenario differently, so it could use a bit of vigilance. Maybe running a simple benchmark in the future would be a standard task when extending LoRA support for new models?

Going back to the topic at hand. @BenjaminBossan @sayakpaul I wouldn't want to overcomplicate this and code churn making you review multiple times, would you be interested in me opening the PR with:

  1. PEFT with universal lossy conversion
  2. PEFT with custom conversion (i.e. how it is now)
  3. Both, so we can run benchmarks and compare directly.

vladmandic commented on Mar 24, 2026

@vladmandic
Contributor

@BenjaminBossan @sayakpaul SD.Next has developed its own quantization engine called SDNQ which is getting quite popular and also getting some use outside of SD.Next itself since its released as a separate no-dependency package.
Highlights are that a) it has 30+ int and fp quant types, both standard and SVD-style, b) little-to-no overhead, c) GPU-agnostic, confirmed working on CUDA/ROCm/Ipex/etc., d) supports quantization load-time / post-load / pre-quantizized, e) user can use standard diffusers/transfomers load/save method to save quantized model and load it as such.
few examples:

So yes, absolutely, we'd love to have peft working with sdnq
Perhaps we should have a separate conversation on this topic?

sayakpaul commented on Mar 24, 2026

@sayakpaul
Member

@vladmandic sure, feel free to open a discussion regarding it in peft. And glad to see the growth of SDNQ.

Maybe running a simple benchmark in the future would be a standard task when extending LoRA support for new models?

Not sure I follow. It only comes to picture when X-to-LoRA conversion is concerned from what I understand. I think running simple benchmarks are fine as long as it doesn't terribly block a new feature integration because LoRAs are popular and probably inseparable at this point.

would you be interested in me opening the PR with:

I think (3) will be quite generous and probably the most meaningful in getting a good first taste of the LoRA conversion method being shipped in peft. I think this will also help us get a sense of how it might pan out for models to come. Thanks very much for offering to help here.

CalamitousFelicitousness commented on Mar 24, 2026

@CalamitousFelicitousness
Contributor

Will action on (3).

Not sure I follow. It only comes to picture when X-to-LoRA

Yeah, my brain used "LoRA" as shorthand for "Adapter" there.

BenjaminBossan commented on Mar 25, 2026

@BenjaminBossan
Member

I agree that going with 3 makes sense here, especially since the work is already done. LoKr is a bit of a special case in this ecosystem since it has somewhat moderate adoption (though still distant second behind LoRA). For other PEFT methods, writing direct conversions wouldn't make sense and in many cases that isn't even possible.

Regarding benchmarking, I did consider adding a metric to the PEFT image-gen benchmark I mentioned above which would convert the non-LoRA adapter to LoRA and then evaluate it on the test set again, to see how much we lose through conversion. It is not quite as straightforward as we have a free parameter (rank) and sweeping over it is expensive. Also, not all PEFT methods support conversion (yet). But I think it would make sense to add this in the future.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions