-
Notifications
You must be signed in to change notification settings - Fork 729
[2/4] Register each GGML IQ format once for dispatch and export #2525
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -121,6 +121,59 @@ def fake_quantize_with_cache( | |
| return inputs + (reconstructed - inputs).detach() | ||
|
|
||
|
|
||
| @dataclass(frozen=True) | ||
| class IQFormat: | ||
| """Everything backend dispatch and export need to know about one IQ format. | ||
|
|
||
| Each format module declares one of these beside its encoder and decoder, and | ||
| :data:`~modelopt.torch.quantization.ggml.registry.IQ_FORMAT_REGISTRY` lists them. The | ||
| per-format pieces -- codebook, search, payload layout -- stay in the format's module; what | ||
| lives here is the part every format does the same way. | ||
| """ | ||
|
|
||
| name: str | ||
| block_size: int | ||
| block_bytes: int | ||
| quantize: Callable[..., tuple[torch.Tensor, torch.Tensor]] | ||
| dequantize: Callable[..., torch.Tensor] | ||
| # Encode and decode are chunked separately: packing runs once per weight and is bounded by | ||
| # its search temporaries, decoding runs every forward and is bounded by kernel launches. | ||
| block_chunk_size: int | ||
| decode_chunk_size: int | ||
|
|
||
| @property | ||
| def effective_bits(self) -> float: | ||
| """Packed storage cost per weight.""" | ||
| return self.block_bytes * 8 / self.block_size | ||
|
|
||
| def fake_quant( | ||
| self, | ||
| inputs: torch.Tensor, | ||
| quantizer, | ||
| *, | ||
| block_chunk_size: int | None = None, | ||
| decode_chunk_size: int | None = None, | ||
| ) -> torch.Tensor: | ||
| """TensorQuantizer backend for this format, with pass-through backward.""" | ||
| if getattr(quantizer, "num_bits", None) != self.name: | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
This guard replaces three per-format copies of the same check, and I can't find a test that hits it —
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Done in |
||
| raise ValueError( | ||
| f"The ggml {self.name.upper()} backend requires num_bits={self.name!r}" | ||
| ) | ||
| return fake_quantize_with_cache( | ||
| inputs, | ||
| quantizer, | ||
| format_name=self.name, | ||
| block_chunk_size=( | ||
| self.block_chunk_size if block_chunk_size is None else block_chunk_size | ||
| ), | ||
| decode_chunk_size=( | ||
| self.decode_chunk_size if decode_chunk_size is None else decode_chunk_size | ||
| ), | ||
| quantize=self.quantize, | ||
| dequantize=self.dequantize, | ||
| ) | ||
|
|
||
|
|
||
| def narrow_to_float32(blocks: torch.Tensor) -> torch.Tensor: | ||
| """Narrow ``blocks`` to float32 the way the CUDA ``load_float`` helper does. | ||
|
|
||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -38,7 +38,7 @@ | |
| from .codebooks import iq1_s_grid_bytes | ||
| from .common import ( | ||
| GGML_BLOCK_SIZE, | ||
| fake_quantize_with_cache, | ||
| IQFormat, | ||
| narrow_to_float32, | ||
| validate_block_chunk_size, | ||
| validate_packed_weights, | ||
|
|
@@ -235,22 +235,17 @@ def dequantize_iq1_s( | |
| return decoded.reshape(shape) | ||
|
|
||
|
|
||
| def iq1_s_fake_quant( | ||
| inputs: torch.Tensor, | ||
| quantizer, | ||
| *, | ||
| block_chunk_size: int = _DEFAULT_BLOCK_CHUNK_SIZE, | ||
| decode_chunk_size: int = _DEFAULT_DECODE_CHUNK_SIZE, | ||
| ) -> torch.Tensor: | ||
| """IQ1_S weight backend for TensorQuantizer, with pass-through backward.""" | ||
| if getattr(quantizer, "num_bits", None) != "iq1_s": | ||
| raise ValueError("The ggml IQ1_S backend requires num_bits='iq1_s'") | ||
| return fake_quantize_with_cache( | ||
| inputs, | ||
| quantizer, | ||
| format_name="iq1_s", | ||
| block_chunk_size=block_chunk_size, | ||
| decode_chunk_size=decode_chunk_size, | ||
| quantize=quantize_iq1_s, | ||
| dequantize=dequantize_iq1_s, | ||
| ) | ||
| IQ1_S_FORMAT = IQFormat( | ||
| name="iq1_s", | ||
| block_size=IQ1_S_BLOCK_SIZE, | ||
| block_bytes=IQ1_S_BLOCK_BYTES, | ||
| quantize=quantize_iq1_s, | ||
| dequantize=dequantize_iq1_s, | ||
| block_chunk_size=_DEFAULT_BLOCK_CHUNK_SIZE, | ||
| decode_chunk_size=_DEFAULT_DECODE_CHUNK_SIZE, | ||
| ) | ||
|
|
||
| # Kept for callers of the per-format entry point. The record captured quantize_iq1_s and | ||
| # dequantize_iq1_s when it was built, so patching those module functions changes neither backend | ||
| # dispatch nor this alias; substitute a format's encoder or decoder in IQ_FORMAT_REGISTRY. | ||
| iq1_s_fake_quant = IQ1_S_FORMAT.fake_quant | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
Note the alias now captures
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Documented in |
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Deriving
IQ_FORMATSfrom the registry means any format registered for backend dispatch is automatically declared exportable by both exporters and byconvert_hf_config. That is fine today since every record carries a packer and geometry, but it removes the ability to land a QAT-only format ahead of its export path. Worth stating that as intended in the PR body.There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Yes, that's intended.
789ea021cadds a comment aboveIQ_FORMATSinquant_format.pysaying so, and the PR body now has a design-choice bullet. The reason: fake quant isdequantize(quantize(w)), so a format can't be dispatched without the packer and block geometry, and those are all export reads. A QAT-only IQ format can't exist.