Skip to content

fix: update ggml to fix INT8 convrot backend fallback - #2069

Merged
leejet merged 1 commit into
masterfrom
fix/int8-convrot-backend-fallback
Sep 27, 2026
Merged

leejet merged 1 commit into
masterfrom
fix/int8-convrot-backend-fallback

Conversation

@leejet

@leejet leejet commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Summary

Update ggml to preserve INT8 ConvRot packed-input metadata across backend copies. This prevents CPU matmul assertions when activation quantization runs on CUDA and matmul falls back to CPU, as reported on V100.

The fix also unifies packed-input validation across CPU, CUDA, and Vulkan.

Related Issue / Discussion

Fix #2068.

Additional Information

N/A

Checklist

@leejet
leejet merged commit 27f7c43 into master Sep 27, 2026
10 checks passed
@leejet
leejet deleted the fix/int8-convrot-backend-fallback branch September 27, 2026 12:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Qwen Image 2.1 INT8 ConvRot aborts at CPU matmul on CUDA

1 participant