Skip to content

fix: keep scaled INT8 convrot matmuls on Vulkan - #2070

Merged
leejet merged 1 commit into
masterfrom
fix/vulkan-scaled-int8-convrot
Sep 27, 2026
Merged

leejet merged 1 commit into
masterfrom
fix/vulkan-scaled-int8-convrot

Conversation

@leejet

@leejet leejet commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Summary

Explicitly quantize scaled convrot activations before INT8 matmul, preventing CPU fallback in Qwen attention output projections.

Related Issue / Discussion

N/A

Additional Information

N/A

Checklist

@leejet
leejet merged commit 42ab1c1 into master Sep 27, 2026
10 checks passed
@leejet
leejet deleted the fix/vulkan-scaled-int8-convrot branch September 27, 2026 12:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant