[6/6] Add the IQ1_M CUDA encoder and register the format - #2595
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (6)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. 📝 WalkthroughWalkthroughThe change adds IQ1_M as a GGML-compatible weight-only format with CUDA packing and a Python quantization path. It adds PTQ recipe support, registers the format, and extends documentation and tests. ChangesIQ1_M support
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant quantize_iq1_m
participant iq1_m_pack
participant iq1_m_pack_cuda
participant IQ1_M_encoder
quantize_iq1_m->>iq1_m_pack: input, grid, predicted scales
iq1_m_pack->>iq1_m_pack_cuda: contiguous packing tensors
iq1_m_pack_cuda->>IQ1_M_encoder: launch block encoder
IQ1_M_encoder-->>iq1_m_pack_cuda: packed payloads
iq1_m_pack_cuda-->>iq1_m_pack: packed byte tensor
iq1_m_pack-->>quantize_iq1_m: packed weights
Suggested reviewers: Merge Risk: ⚪ Minimal · up to IQ1_M is wired through quantization, export, and recipes, and the inspected CUDA path preserves the expected block layout and scale order. No concrete merge-blocking behavior is indicated; proceed with normal checks. 🚥 Pre-merge checks | ✅ 5 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (5 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #2595 +/- ##
==========================================
+ Coverage 69.46% 78.86% +9.39%
==========================================
Files 611 611
Lines 68205 68219 +14
==========================================
+ Hits 47380 53798 +6418
+ Misses 20825 14421 -6404
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
5472160 to
e492750
Compare
### What does this PR do? Type of change: new feature (not yet user-reachable) **First of two PRs adding IQ1_M** at 1.75 bits per weight, just above IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is deliberately **not registered**, so no quantizer dispatches to it and the `ggml` package does not export it. #2595 adds the CUDA encoder, registers the format and adds its recipe. With both, ModelOpt supports all five GGML IQ formats at one and two bits. On the mixed-precision checkpoint #2511 measured (`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B parameters**. With all five formats we can read 89.0% of that file; the rest is k-quants and F32. ### What's distinctive about it **IQ1_M is the most irregular layout of the five.** There is no leading block scale field at all. The FP16 super-block scale is reassembled from the top nibble of each of four scale words: ```c scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000); ``` It is also finer grained than IQ1_S: a local scale per **two** groups rather than four, and a delta shift chosen **per group** rather than per sub-block. That is where its extra 0.1875 bits go. ### Shared with IQ1_S rather than copied IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same ±1/8 delta, every (shift, local scale) choice for every 8-value vector. It differs only in how it selects among those choices afterwards. So the search moves out of IQ1_S's encoder into `_search_shifted_grid`, which both call, and `iq1_m.py` keeps only its selection and packing. **IQ1_S's encoded bytes are unchanged**, checked by hashing its output before and after on a fixed input. ### A scale-anchor correction IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a block's peak-to-RMS** rather than being flat, and clamps higher. It uses `clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat `0.61`. Measured over 15 Qwen3.8-27B MLP weights: | | flat 0.61 | correct anchor | | |---|---|---|---| | relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** | It is consistent on every tensor, with no outliers. The anchor changes quality without touching layout, so neither round-trip nor conformance tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms` now pins it, for all five formats; see Testing. ### Family parity Two surface asymmetries close here, so the five are uniform. `IQ1_S` now exposes `_predict_iq1_s_scales` like the other four, instead of computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the IQ1_S table it shares. ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped:** ``` IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0 ``` This mattered: **my first IQ1_M decoder had a real bug.** A `repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder still passed, because the encoder made the matching mistake. Only comparison against bytes we did not produce caught it. Blocks from that checkpoint ship as conformance vectors, and mutation testing confirms they catch a mis-set scale nibble. The decoder unpacks every field in one vectorized pass, since fake quant decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3). - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M codec cases, including the llama.cpp conformance check - `test_scale_anchor_follows_peak_to_rms` pins every format's scale anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4, 8 and 16, reaching both clamps and two points on each slope, and compares them against anchors written out in the test. Mutations each fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61, moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's CUDA-vs-PyTorch parity still holds after its encoder refactor. - IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash identically to the pre-split version of this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds no codebook; it reuses the IQ1_S table already carried in `codebooks.py`. The new conformance vectors come from `unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model `Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595 carries the entry. - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration), all merged → **this** → #2595 (IQ1_M CUDA encoder and registration). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ1_M quantization and dequantization for compact, GGML-compatible blocks of 256 values. * Added access to the IQ1_M grid and configurable chunk sizes for processing data. * **Bug Fixes** * Improved IQ1_S scale prediction and grid-search organization while preserving its encoding behavior. * **Tests** * Added IQ1_M conformance data and included the format in shared IQ-format test coverage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Second of two changes adding IQ1_M. The previous change landed the PyTorch codec; this one adds its CUDA encoder and makes the format reachable. In the kernel the delta shift is free per group, so it sits above the entry index in the sort key: a tie still prefers the lower shift and then the lower entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels read it from global memory and rely on the cache. On a 5632x2048 weight the encoder runs at 318 M elem/s against the torch search's 5.6, and its packed bytes match the PyTorch encoder's. The two IQ1 kernels load each vector, score it against a grid entry and apply the +/-1/8 shift the same way, so those three steps move into common.cuh as load_vector, grid_terms and shifted_error, and IQ1_S uses them too. IQ1_S's packed bytes are unchanged. IQ1_M gets an IQFormat record and one IQ_FORMAT_REGISTRY entry, so backend dispatch, both exporters and convert_hf_config take it from there. The ggml package exports it, and a general/ptq/iq1_m recipe uses it. Registering it brings it under every registry-driven test with no IQ1_M-specific test code, and the hf_ptq example test gains an IQ1_M case. IQ1_M covers 25 tensors and 1.2B parameters of the mixed-precision checkpoint IQ2_XXS measured. With it, ModelOpt supports all five GGML IQ formats at one and two bits. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
e492750 to
49f8d1e
Compare
meenchen
left a comment
There was a problem hiding this comment.
Bot review (gpt-6-astra) — DM the bot to share feedback.
LGTM on correctness and coverage; human sign-off remains for the GGML-referenced layout and justified extensions to existing tests.
Needs action:
- Review and sign off on licensing for the GGML layout reference in
modelopt/torch/kernels/quantization/ggml/iq1_m.cu; its NVIDIA header matchesLICENSE_HEADER, andLICENSEalready includes GGML’s MIT notice. - Sign off on the existing CUDA, recipe, and example test extensions: they add IQ1_M cases without removing or weakening assertions.
No action needed:
- Design review: the problem is accelerating and exposing the existing IQ1_M codec. Compared with PyTorch-only encoding and duplicating IQ1_S machinery, the reported speedup and shared helper extraction justify this approach. It extends the existing registry and Torch/CUDA extension rather than introducing another subsystem.
- Checked nibble packing, distributed FP16 scale bits, shift/index tie ordering, zero handling, synchronization, fallback, and registry-driven dispatch/export coverage; no correctness issues found.
- Tests were inspected, not executed in this review.
### What does this PR do? Type of change: documentation Changes the PR sizing rule in `AGENTS.md` to count only **added source** lines toward the ~500-line budget, instead of total changed lines. Deletions are cheap to review, so a PR that mostly removes code shouldn't be pushed into a split. Tests and docs are excluded too, since every sub-PR has to carry its own tests. The check uses the insertions count from `git diff --shortstat` with a pathspec that excludes `tests/` and `docs/`. ### Usage ```bash git diff --shortstat origin/main...HEAD -- . ':!tests' ':!docs' # N files changed, X insertions(+), Y deletions(-) -> compare X against ~500 ``` ### Testing - `pre-commit run --files AGENTS.md` (markdownlint passes). - Ran the pathspec against recent commits (#2595, #2513) to confirm it drops test and doc lines from the count. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Follow-up to #2494, which introduced the sizing guidance. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated review guidance to measure pull request size by added source lines, excluding deletions, tests, and documentation. The guidance retains the recommendation to check the size before opening a review. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
### What does this PR do? Type of change: performance Two fixes that make fake-quantized GGML IQ models fast. Both were found by instrumenting the hf_ptq example (TinyLlama, 154 IQ weights, 100-token preview) and counting every encode and decode per weight. 1. **Each weight is packed once.** Fake quant packs a weight on its first forward and caches the payload, but export then ran the search again on the same unchanged weight: 154 more encodes for 154 weights, in all five formats. `IQFormat.pack` now reuses the cached payload when it was packed from the weight being exported, so the checkpoint also holds exactly the bytes the evaluated model decoded. 2. **Decoding runs on CUDA.** All five formats had CUDA encoders but decoded with PyTorch ops. Fake quant decodes every weight on every forward, so once packing was fast and cached, decoding was **55–66% of the run**, about 3 ms per weight per forward; on a 27B model that is about 10 s per forward. Each format's `Format` struct from #2615 gains a bit-exact `decode()` beside its `store()`. ### Results The instrumented hf_ptq example on the same GPU, with the extension already built: | format | total wall before → after | encodes at export | decode time before → after | |---|---|---|---| | IQ1_S | 55.6 → **21.9 s** | 154 → **0** | 32.0 → **1.4 s** | | IQ1_M | 69.8 → **21.0 s** | 154 → **0** | 46.3 → **1.3 s** | | IQ2_XXS | 61.6 → **19.0 s** | 154 → **0** | 41.9 → **1.3 s** | | IQ2_XS | 63.8 → **19.1 s** | 154 → **0** | 44.3 → **1.3 s** | | IQ2_S | 55.6 → **19.9 s** | 154 → **0** | 35.8 → **1.3 s** | Fake quant still packs each weight exactly once and decodes it 100 times, once per preview token. IQ1_M was the slowest format before because its PyTorch decoder did the most work; it now matches IQ1_S. Decoding one 5632×2048 weight goes from 3.3–5.2 ms to **0.06–0.14 ms**. ### How **Reusing the payload.** TensorQuantizer hands fake quant the weight reshaped into 256-value blocks, so the cache is keyed on that view, while export holds the weight itself. An exact-key match therefore never hit in a real model, even though a 256-wide unit test passed. `pack` also accepts another contiguous view of the same storage, version and length: the same values in the same order, hence the same GGML blocks. It then reshapes the payload to the weight's layout. The version counter still makes a weight edited after its forward pack afresh. The Megatron exporter keeps packing on its own, since it can remap or slice a weight before packing it. **Bit-exact decoders.** One CUDA thread decodes one 8-value vector. Every float operation is explicitly rounded (`__fmul_rn`, `__fadd_rn`, `__fdiv_rn`) in the PyTorch decoder's order, so the compiler cannot fuse a multiply into an add, and the output is bit-identical to the PyTorch decoders. `dequantize_<format>` uses the unpacker for CUDA payloads and keeps the PyTorch path otherwise. ### Testing - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **74 passed** on RTX PRO 6000 (sm_120). That includes 25 new decoder cases: for each format, CUDA equals the PyTorch decoder bit for bit in four dtypes, on random payloads (including block scales that decode to inf or NaN) and real encodings. The CUDA path also reproduces llama.cpp's values on the conformance blocks. Dropping IQ2_XXS's parity bit in the kernel fails all five IQ2_XXS cases. - `tests/unit/torch/export/test_export_weight.py`: export reuses the cached payload without calling the encoder, and repacks a weight edited after its forward, for all five formats. The weight is 512 wide so the blocked view really differs; with an exact-key match only, the reuse test fails for every format. The IQ payload export test covers all five formats rather than two. - IQ unit tests (`test_ggml_backend.py`, `test_iq_formats.py`, `test_convert_hf_config.py`, `test_presets.py`, `test_export_weight.py`): **186 passed** - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml'`: **45 passed** in `nvcr.io/nvidia/nemo:26.08` - `tests/examples/hf_ptq/test_llm_ptq.py -k iq`: **5 passed**, 27.7–30.9 s each - On this GPU, `test_torch_extensions.py` still shows #2515's two Q8_0 failures, which fail identically on a clean `main`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Same bindings, error messages, exported bytes and decoded values, only faster. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A. It speeds up formats added in this unreleased cycle without changing their output. - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Builds on #2615 (one CUDA encoder per IQ family), now merged. Follows the IQ series (#2511, #2525, #2512, #2565, #2513, #2595), all merged. The benchmark above was measured on the combined branch before the split. After rebasing onto `main` with #2615 merged, this PR's code is the same apart from `kVectorsPerBlock` moving into `common.cuh`'s shared constants and #2615's validation-order fix. On the rebased branch, all 40 encoder and decoder hashes match, 74 GPU tests pass and 186 unit tests pass. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
### What does this PR do? Type of change: new feature Adds Q8_0 encoding, decoding, and the GGML fake-quant backend. Each 32-weight block stores one FP16 scale and 32 signed int8 values in 34 bytes (8.5 bits per weight). - Generalizes dispatch to `GGML_FORMAT_REGISTRY` / `GGMLFormat` while preserving `IQFormat` as a type alias. `IQ_FORMAT_REGISTRY` is an IQ-only compatibility dictionary sharing the same format records, not a mutation-propagating view. - Retains IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, and IQ2_S registrations alongside Q8_0. - Uses the merged `q8_0_pack` CUDA extension when available and the PyTorch encoder otherwise. - Matches canonical reciprocal-then-multiply rounding, using the unrounded FP32 scale to choose int8 values and FP16 only for serialized scale storage. - Adds exact-byte rounding regressions and checks that the compatibility registry shares the same format records. Removes a redundant zero-block buffer copy. Checkpoint export, recipes, documentation, and the user-facing Q8_0 changelog remain in #2517. ### Base and dependencies The target remains `main`. Current head `71e3183a7b7eac7e5e3888ee45a206c572028af9` includes main at `67a68f8fd4902a7b67c78a2f00f40562b081fe72`, including #2595 (IQ1_M registration), #2615 (shared IQ CUDA encoders), and #2604 (packed-weight cache reuse and CUDA IQ decoding). All five IQ formats, Q8_0, and the compatibility aliases are retained. The comparison against `main` contains only the intended 12 Q8_0 files, with 207 added source lines excluding tests and docs. The inherited IQ refactor and export changes are not part of this PR's diff. #2515 has already merged and supplies the Q8_0 kernel. This PR retains the small reciprocal-rounding correction to that kernel needed for exact reference parity. ### Usage ```python import torch from modelopt.torch.quantization.ggml import quantize_q8_0, dequantize_q8_0 weight = torch.randn(2, 64, dtype=torch.bfloat16) packed, shape = quantize_q8_0(weight) restored = dequantize_q8_0(packed, shape) ``` The final weight dimension must be divisible by 32. Backend dispatch also accepts `num_bits="q8_0"`, `backend="ggml"`, and `block_sizes={-1: 32}`. ### Testing - Current head `71e3183a7`: **183 focused CPU tests passed**, covering Q8_0, all six registered backends, IQ formats, export metadata, and recipe presets. - Current-head applicable pre-commit checks passed: Ruff, formatting, mypy, CUDA formatting, license headers, security checks, merge markers, line endings, and file size. - Previous head `2beb70ffa`: **all six Q8_0 CUDA cases passed**, including exact-byte reciprocal rounding and unrounded-scale regressions; [GPU job log](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36890153545/job/110468900810). That GPU lane completed with 1,663 passed and 67 skipped. The overall workflow was cancelled after a different lane was cancelled; it is not an all-green workflow result. - Current head `71e3183a7`: the GPU CI mirror now points to the exact PR head. [GPU CI](https://github.com/NVIDIA/Model-Optimizer/actions/runs/37061248628), [example CI](https://github.com/NVIDIA/Model-Optimizer/actions/runs/37061248555), and [regression CI](https://github.com/NVIDIA/Model-Optimizer/actions/runs/37061248637) are running. Previous-head results do not validate this new head. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: yes; existing IQ names and registrations are retained. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`?: yes; no new dependency. - Did you write any new necessary tests?: yes. - Did you update `CHANGELOG.rst`?: N/A here; #2517 carries the single Q8_0 feature entry. - Did you get Claude approval on this PR?: pending current-head review. ### Related PRs 1. [#2515 — Q8_0 CUDA packing kernel](#2515) — merged. 2. [#2595 — IQ1_M registration](#2595) — merged. 3. **#2516 — Q8_0 quantization codec and backend** — this PR. 4. [#2517 — Q8_0 checkpoint export and recipes](#2517) — follows this PR. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Q8_0 quantization support, including weight packing, unpacking, and fake quantization. * Added Q8_0 to the available GGML formats, alongside existing IQ formats, through a shared quantization interface. * Added support for formats with different block sizes when validating weights. * Q8_0 uses CUDA acceleration when available and falls back to PyTorch when needed. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
What does this PR do?
Type of change: new feature
Second of two PRs adding IQ1_M (1.75 bits per weight). #2513 landed the PyTorch codec; this PR adds its CUDA encoder and makes the format reachable. With it, ModelOpt supports all five GGML IQ formats at one and two bits.
quantize_iq1_mIQFormatrecord and oneIQ_FORMAT_REGISTRYentry, so backend dispatch, both exporters andconvert_hf_configtake it from thereggmlpackage exportgeneral/ptq/iq1_mrecipe, its presets,ptq.mdand a CHANGELOG entryThe kernel lands with the registration so every registered format keeps a CUDA encoder.
The kernel
In the kernel the delta shift is free per group, so it sits above the entry index in the sort key: a tie still prefers the lower shift and then the lower entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels read it from global memory and rely on the cache.
Shared with IQ1_S rather than copied
The two IQ1 kernels load each vector, score it against a grid entry and apply the ±1/8 shift the same way. So those three steps move into
common.cuhasload_vector,grid_termsandshifted_error, and IQ1_S uses them too. IQ1_S's packed bytes are unchanged: its CUDA output on a 5632×2048 weight hashes the same before and after, and so does IQ1_M's, compared against the pre-split version of this change. IQ1_S encodes at the same speed (306 M elem/s).Usage
Testing
Registering the format brings it under every registry-driven test with no IQ1_M-specific test code: backend dispatch, weight caching, the
num_bitsguard,convert_hf_configmetadata, Megatron export and theTensorQuantizertests in the shared battery. The shared CUDA battery gains one row.tests/unit/torch/quantization/test_ggml_backend.py,test_iq_formats.py,tests/unit/torch/export/test_convert_hf_config.py,tests/unit/recipe/test_presets.py: 166 passed-k 'ggml or iq or gguf or registry'over quantization, export and recipe tests): 221 passed. The one failure,test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes, is atorchvisionimport error in my environment, unrelated to IQ.tests/gpu/torch/quantization/test_iq_formats_cuda.py,test_iq1_s_cuda.py,test_iq2_xs_cuda.py: 49 passed on RTX PRO 6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch encoder paritytests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml': 45 passed (9 tests × 5 formats) innvcr.io/nvidia/nemo:26.08tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m: passedgeneral/ptqnow holds 31 recipes;ptq.mdis updated.Rebased onto
mainafter #2513 merged. The resulting tree is identical to the one the runs above tested, and the unit set was rerun on it: 166 passed.On this GPU, two of #2515's Q8_0 tests in
tests/gpu/_extensions/test_torch_extensions.pyfail:test_cuda_ext_q8_0_zero_and_roundf_layoutandtest_cuda_ext_q8_0_dequantizes_with_small_error. They fail identically on a cleanmaincheckout, so they are not from this PR.Before your PR is "Ready for review"
CONTRIBUTING.md: ✅ No new code sources or dependencies.Additional Information
Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M codec), all merged → this.
🤖 Generated with Claude Code
Summary by CodeRabbit