Skip to content

fix: load INT8 convrot LLM embeddings correctly - #2067

Merged
leejet merged 1 commit into
masterfrom
fix/llm-int8-convrot-embeddings
Sep 27, 2026
Merged

leejet merged 1 commit into
masterfrom
fix/llm-int8-convrot-embeddings

Conversation

@leejet

@leejet leejet commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • Exclude quantization scale tensors from LLM and vision shape detection, preventing incorrect dimensions and initialization failures.
  • Add CPU lookup for INT8 embeddings with scalar or per-token scales and optional inverse convrot.

Related Issue / Discussion

N/A

Additional Information

N/A

Checklist

@leejet
leejet merged commit ede32a6 into master Sep 27, 2026
10 checks passed
@leejet
leejet deleted the fix/llm-int8-convrot-embeddings branch September 27, 2026 12:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant