Skip to content

[6/6] Add the IQ1_M CUDA encoder and register the format - #2595

Merged
cjluo-nv merged 1 commit into
mainfrom
chenjiel/iq1-m-register
Oct 1, 2026
Merged

cjluo-nv merged 1 commit into
mainfrom
chenjiel/iq1-m-register

Conversation

@cjluo-nv

@cjluo-nv cjluo-nv commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Type of change: new feature

Second of two PRs adding IQ1_M (1.75 bits per weight). #2513 landed the PyTorch codec; this PR adds its CUDA encoder and makes the format reachable. With it, ModelOpt supports all five GGML IQ formats at one and two bits.

  • the CUDA encoder, its binding and extension build wiring, plus the CUDA path in quantize_iq1_m
  • an IQFormat record and one IQ_FORMAT_REGISTRY entry, so backend dispatch, both exporters and convert_hf_config take it from there
  • the ggml package export
  • the general/ptq/iq1_m recipe, its presets, ptq.md and a CHANGELOG entry

The kernel lands with the registration so every registered format keeps a CUDA encoder.

The kernel

In the kernel the delta shift is free per group, so it sits above the entry index in the sort key: a tie still prefers the lower shift and then the lower entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels read it from global memory and rely on the cache.

5632×2048 weight torch CUDA
IQ1_M encode 5.6 M elem/s 318 M elem/s 57×

Shared with IQ1_S rather than copied

The two IQ1 kernels load each vector, score it against a grid entry and apply the ±1/8 shift the same way. So those three steps move into common.cuh as load_vector, grid_terms and shifted_error, and IQ1_S uses them too. IQ1_S's packed bytes are unchanged: its CUDA output on a 5632×2048 weight hashes the same before and after, and so does IQ1_M's, compared against the pre-split version of this change. IQ1_S encodes at the same speed (306 M elem/s).

Usage

python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m

Testing

Registering the format brings it under every registry-driven test with no IQ1_M-specific test code: backend dispatch, weight caching, the num_bits guard, convert_hf_config metadata, Megatron export and the TensorQuantizer tests in the shared battery. The shared CUDA battery gains one row.

  • tests/unit/torch/quantization/test_ggml_backend.py, test_iq_formats.py, tests/unit/torch/export/test_convert_hf_config.py, tests/unit/recipe/test_presets.py: 166 passed
  • broader unit sweep (-k 'ggml or iq or gguf or registry' over quantization, export and recipe tests): 221 passed. The one failure, test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes, is a torchvision import error in my environment, unrelated to IQ.
  • tests/gpu/torch/quantization/test_iq_formats_cuda.py, test_iq1_s_cuda.py, test_iq2_xs_cuda.py: 49 passed on RTX PRO 6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch encoder parity
  • tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml': 45 passed (9 tests × 5 formats) in nvcr.io/nvidia/nemo:26.08
  • tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m: passed
  • reconstruction error falls monotonically across all five formats, pinned by a test
  • general/ptq now holds 31 recipes; ptq.md is updated.

Rebased onto main after #2513 merged. The resulting tree is identical to the one the runs above tested, and the unit set was rerun on it: 166 passed.

On this GPU, two of #2515's Q8_0 tests in tests/gpu/_extensions/test_torch_extensions.py fail: test_cuda_ext_q8_0_zero_and_roundf_layout and test_cuda_ext_q8_0_dequantizes_with_small_error. They fail identically on a clean main checkout, so they are not from this PR.

Before your PR is "Ready for review"

  • Is this change backward compatible?: ✅
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: ✅ No new code sources or dependencies.
  • Did you write any new necessary tests?: ✅
  • Did you update Changelog?: ✅
  • Did you get Claude approval on this PR?: ❌ Not yet run.

Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M codec), all merged → this.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Added IQ1_M weight-only quantization at 1.75 bits per weight, with CUDA acceleration and a 256-value block size.
    • Added an IQ1_M post-training quantization recipe for eligible linear layers; calibration data is not required.
    • Added IQ1_M to the supported GGML-compatible formats and recipe listings.

@cjluo-nv
cjluo-nv requested review from a team as code owners September 29, 2026 22:53
@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 1e422b5c-294f-469d-8063-cf79829e98c3

📥 Commits

Reviewing files that changed from the base of the PR and between e492750 and 49f8d1e.

📒 Files selected for processing (6)
  • modelopt/torch/export/quant_format.py
  • modelopt/torch/quantization/ggml/__init__.py
  • modelopt/torch/quantization/ggml/iq1_m.py
  • modelopt/torch/quantization/ggml/registry.py
  • tests/gpu/torch/quantization/test_iq_formats_cuda.py
  • tests/unit/recipe/test_presets.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Walkthrough

Walkthrough

The change adds IQ1_M as a GGML-compatible weight-only format with CUDA packing and a Python quantization path. It adds PTQ recipe support, registers the format, and extends documentation and tests.

Changes

IQ1_M support

Layer / File(s) Summary
Shared IQ1 scoring helpers
modelopt/torch/kernels/quantization/ggml/common.cuh, modelopt/torch/kernels/quantization/ggml/iq1_s.cu
Shared helpers calculate vector statistics and shifted quantization error. IQ1_S uses these helpers in its scoring loops.
IQ1_M CUDA packing
modelopt/torch/kernels/quantization/ggml/iq1_m.cu, modelopt/torch/kernels/quantization/ggml/ggml.cpp, modelopt/torch/quantization/extensions.py, tests/gpu/torch/quantization/test_iq_formats_cuda.py
The CUDA encoder packs 256-value blocks into 56-byte payloads. The C++ binding validates inputs, the extension builds the encoder, and CUDA format tests include IQ1_M.
Python format and registration
modelopt/torch/quantization/ggml/iq1_m.py, modelopt/torch/quantization/ggml/__init__.py, modelopt/torch/quantization/ggml/registry.py, modelopt/torch/export/quant_format.py
The Python quantizer uses the CUDA packer when the input is on CUDA and the GGML extension is available. IQ1_M format metadata, exports, registry membership, and a public format constant are added.
PTQ recipes and validation
modelopt_recipes/configs/numerics/iq1_m.yaml, modelopt_recipes/configs/ptq/presets/model/iq1_m.yaml, modelopt_recipes/general/ptq/iq1_m.yaml, modelopt_recipes/ptq.md, tests/unit/recipe/test_presets.py, tests/examples/hf_ptq/test_llm_ptq.py, CHANGELOG.rst
The IQ1_M numeric configuration, weight-only preset, and PTQ recipe are added. Documentation and recipe tests include IQ1_M at 1.75 bits per weight.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant quantize_iq1_m
  participant iq1_m_pack
  participant iq1_m_pack_cuda
  participant IQ1_M_encoder
  quantize_iq1_m->>iq1_m_pack: input, grid, predicted scales
  iq1_m_pack->>iq1_m_pack_cuda: contiguous packing tensors
  iq1_m_pack_cuda->>IQ1_M_encoder: launch block encoder
  IQ1_M_encoder-->>iq1_m_pack_cuda: packed payloads
  iq1_m_pack_cuda-->>iq1_m_pack: packed byte tensor
  iq1_m_pack-->>quantize_iq1_m: packed weights
Loading

Suggested reviewers: hychiang-git

Merge Risk: ⚪ Minimal · up to 49f8d

IQ1_M is wired through quantization, export, and recipes, and the inspected CUDA path preserves the expected block layout and scale order. No concrete merge-blocking behavior is indicated; proceed with normal checks.

🚥 Pre-merge checks | ✅ 5 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 61.90% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 21 functions across 14 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Security Anti-Patterns ✅ Passed PASS. The pull request changes five Python files under modelopt/ and no Python files under examples/. The added Python lines contain no torch.load(..., weights_only=False), `numpy.load(..., allo…
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main changes: adding the IQ1_M CUDA encoder and registering the format.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-01 05:24 UTC

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 78.86%. Comparing base (e5b6332) to head (49f8d1e).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2595      +/-   ##
==========================================
+ Coverage   69.46%   78.86%   +9.39%     
==========================================
  Files         611      611              
  Lines       68205    68219      +14     
==========================================
+ Hits        47380    53798    +6418     
+ Misses      20825    14421    -6404     
Flag Coverage Δ
examples-diffusers 21.38% <57.14%> (+0.01%) ⬆️
examples-gpt-oss 13.61% <57.14%> (+0.02%) ⬆️
examples-hf_ptq 23.30% <100.00%> (+0.04%) ⬆️
examples-llm_distill 13.67% <57.14%> (+0.02%) ⬆️
examples-llm_eval 17.62% <57.14%> (+0.02%) ⬆️
examples-llm_qat 17.79% <57.14%> (+0.01%) ⬆️
examples-llm_sparsity 16.06% <57.14%> (+0.02%) ⬆️
examples-megatron_bridge 26.61% <57.14%> (-0.08%) ⬇️
examples-specdec_bench 13.38% <57.14%> (+0.02%) ⬆️
examples-speculative_decoding 17.99% <57.14%> (-0.05%) ⬇️
examples-torch_onnx 21.93% <57.14%> (+0.01%) ⬆️
examples-torch_trt 15.42% <57.14%> (+0.02%) ⬆️
examples-vllm_serve 13.86% <57.14%> (+0.02%) ⬆️
gpu 58.52% <100.00%> (+36.95%) ⬆️
regression 15.30% <57.14%> (+0.02%) ⬆️
unit 59.16% <64.28%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@cjluo-nv
cjluo-nv force-pushed the chenjiel/iq1-m-register branch from 5472160 to e492750 Compare September 29, 2026 23:12
cjluo-nv added a commit that referenced this pull request Sep 30, 2026
### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ1_M** at 1.75 bits per weight, just above
IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2595 adds the CUDA encoder,
registers the format and adds its recipe. With both, ModelOpt supports
all five GGML IQ formats at one and two bits.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B
parameters**. With all five formats we can read 89.0% of that file; the
rest is k-quants and F32.

### What's distinctive about it

**IQ1_M is the most irregular layout of the five.** There is no leading
block scale field at all. The FP16 super-block scale is reassembled from
the top nibble of each of four scale words:

```c
scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000);
```

It is also finer grained than IQ1_S: a local scale per **two** groups
rather than four, and a delta shift chosen **per group** rather than per
sub-block. That is where its extra 0.1875 bits go.

### Shared with IQ1_S rather than copied

IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same
±1/8 delta, every (shift, local scale) choice for every 8-value vector.
It differs only in how it selects among those choices afterwards. So the
search moves out of IQ1_S's encoder into `_search_shifted_grid`, which
both call, and `iq1_m.py` keeps only its selection and packing.
**IQ1_S's encoded bytes are unchanged**, checked by hashing its output
before and after on a fixed input.

### A scale-anchor correction

IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a
block's peak-to-RMS** rather than being flat, and clamps higher. It uses
`clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat
`0.61`. Measured over 15 Qwen3.8-27B MLP weights:

| | flat 0.61 | correct anchor | |
|---|---|---|---|
| relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** |

It is consistent on every tensor, with no outliers. The anchor changes
quality without touching layout, so neither round-trip nor conformance
tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms`
now pins it, for all five formats; see Testing.

### Family parity

Two surface asymmetries close here, so the five are uniform. `IQ1_S` now
exposes `_predict_iq1_s_scales` like the other four, instead of
computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the
IQ1_S table it shares.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0
```

This mattered: **my first IQ1_M decoder had a real bug.** A
`repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where
llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder
still passed, because the encoder made the matching mistake. Only
comparison against bytes we did not produce caught it. Blocks from that
checkpoint ship as conformance vectors, and mutation testing confirms
they catch a mis-set scale nibble.

The decoder unpacks every field in one vectorized pass, since fake quant
decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3).

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M
codec cases, including the llama.cpp conformance check
- `test_scale_anchor_follows_peak_to_rms` pins every format's scale
anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4,
8 and 16, reaching both clamps and two points on each slope, and
compares them against anchors written out in the test. Mutations each
fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61,
moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S
or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's
CUDA-vs-PyTorch parity still holds after its encoder refactor.
- IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash
identically to the pre-split version of this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds
no codebook; it reuses the IQ1_S table already carried in
`codebooks.py`. The new conformance vectors come from
`unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model
`Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration), all merged →
**this** → #2595 (IQ1_M CUDA encoder and registration).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M quantization and dequantization for compact,
GGML-compatible blocks of 256 values.
* Added access to the IQ1_M grid and configurable chunk sizes for
processing data.

* **Bug Fixes**
* Improved IQ1_S scale prediction and grid-search organization while
preserving its encoding behavior.

* **Tests**
* Added IQ1_M conformance data and included the format in shared
IQ-format test coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Second of two changes adding IQ1_M. The previous change landed the PyTorch
codec; this one adds its CUDA encoder and makes the format reachable.

In the kernel the delta shift is free per group, so it sits above the entry
index in the sort key: a tie still prefers the lower shift and then the lower
entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with
IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels
read it from global memory and rely on the cache. On a 5632x2048 weight the
encoder runs at 318 M elem/s against the torch search's 5.6, and its packed
bytes match the PyTorch encoder's.

The two IQ1 kernels load each vector, score it against a grid entry and apply
the +/-1/8 shift the same way, so those three steps move into common.cuh as
load_vector, grid_terms and shifted_error, and IQ1_S uses them too. IQ1_S's
packed bytes are unchanged.

IQ1_M gets an IQFormat record and one IQ_FORMAT_REGISTRY entry, so backend
dispatch, both exporters and convert_hf_config take it from there. The ggml
package exports it, and a general/ptq/iq1_m recipe uses it. Registering it
brings it under every registry-driven test with no IQ1_M-specific test code,
and the hf_ptq example test gains an IQ1_M case.

IQ1_M covers 25 tensors and 1.2B parameters of the mixed-precision checkpoint
IQ2_XXS measured. With it, ModelOpt supports all five GGML IQ formats at one
and two bits.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
@cjluo-nv
cjluo-nv force-pushed the chenjiel/iq1-m-register branch from e492750 to 49f8d1e Compare September 30, 2026 16:48

@meenchen meenchen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bot review (gpt-6-astra) — DM the bot to share feedback.

LGTM on correctness and coverage; human sign-off remains for the GGML-referenced layout and justified extensions to existing tests.

Needs action:

  • Review and sign off on licensing for the GGML layout reference in modelopt/torch/kernels/quantization/ggml/iq1_m.cu; its NVIDIA header matches LICENSE_HEADER, and LICENSE already includes GGML’s MIT notice.
  • Sign off on the existing CUDA, recipe, and example test extensions: they add IQ1_M cases without removing or weakening assertions.

No action needed:

  • Design review: the problem is accelerating and exposing the existing IQ1_M codec. Compared with PyTorch-only encoding and duplicating IQ1_S machinery, the reported speedup and shared helper extraction justify this approach. It extends the existing registry and Torch/CUDA extension rather than introducing another subsystem.
  • Checked nibble packing, distributed FP16 scale bits, shift/index tie ordering, zero handling, synchronization, fallback, and registry-driven dispatch/export coverage; no correctness issues found.
  • Tests were inspected, not executed in this review.

@cjluo-nv
cjluo-nv merged commit aa89722 into main Oct 1, 2026
104 of 118 checks passed
@cjluo-nv
cjluo-nv deleted the chenjiel/iq1-m-register branch October 1, 2026 05:23
cjluo-nv added a commit that referenced this pull request Oct 1, 2026
### What does this PR do?

Type of change: documentation

Changes the PR sizing rule in `AGENTS.md` to count only **added source**
lines toward the ~500-line budget, instead of total changed lines.
Deletions are cheap to review, so a PR that mostly removes code
shouldn't be pushed into a split. Tests and docs are excluded too, since
every sub-PR has to carry its own tests. The check uses the insertions
count from `git diff --shortstat` with a pathspec that excludes `tests/`
and `docs/`.

### Usage

```bash
git diff --shortstat origin/main...HEAD -- . ':!tests' ':!docs'
# N files changed, X insertions(+), Y deletions(-)  -> compare X against ~500
```

### Testing

- `pre-commit run --files AGENTS.md` (markdownlint passes).
- Ran the pathspec against recent commits (#2595, #2513) to confirm it
drops test and doc lines from the count.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Follow-up to #2494, which introduced the sizing guidance.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated review guidance to measure pull request size by added source
lines, excluding deletions, tests, and documentation. The guidance
retains the recommendation to check the size before opening a review.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
cjluo-nv added a commit that referenced this pull request Oct 2, 2026
### What does this PR do?

Type of change: performance

Two fixes that make fake-quantized GGML IQ models fast. Both were found
by instrumenting the hf_ptq example (TinyLlama, 154 IQ weights,
100-token preview) and counting every encode and decode per weight.

1. **Each weight is packed once.** Fake quant packs a weight on its
first forward and caches the payload, but export then ran the search
again on the same unchanged weight: 154 more encodes for 154 weights, in
all five formats. `IQFormat.pack` now reuses the cached payload when it
was packed from the weight being exported, so the checkpoint also holds
exactly the bytes the evaluated model decoded.
2. **Decoding runs on CUDA.** All five formats had CUDA encoders but
decoded with PyTorch ops. Fake quant decodes every weight on every
forward, so once packing was fast and cached, decoding was **55–66% of
the run**, about 3 ms per weight per forward; on a 27B model that is
about 10 s per forward. Each format's `Format` struct from #2615 gains a
bit-exact `decode()` beside its `store()`.

### Results

The instrumented hf_ptq example on the same GPU, with the extension
already built:

| format | total wall before → after | encodes at export | decode time
before → after |
|---|---|---|---|
| IQ1_S | 55.6 → **21.9 s** | 154 → **0** | 32.0 → **1.4 s** |
| IQ1_M | 69.8 → **21.0 s** | 154 → **0** | 46.3 → **1.3 s** |
| IQ2_XXS | 61.6 → **19.0 s** | 154 → **0** | 41.9 → **1.3 s** |
| IQ2_XS | 63.8 → **19.1 s** | 154 → **0** | 44.3 → **1.3 s** |
| IQ2_S | 55.6 → **19.9 s** | 154 → **0** | 35.8 → **1.3 s** |

Fake quant still packs each weight exactly once and decodes it 100
times, once per preview token. IQ1_M was the slowest format before
because its PyTorch decoder did the most work; it now matches IQ1_S.
Decoding one 5632×2048 weight goes from 3.3–5.2 ms to **0.06–0.14 ms**.

### How

**Reusing the payload.** TensorQuantizer hands fake quant the weight
reshaped into 256-value blocks, so the cache is keyed on that view,
while export holds the weight itself. An exact-key match therefore never
hit in a real model, even though a 256-wide unit test passed. `pack`
also accepts another contiguous view of the same storage, version and
length: the same values in the same order, hence the same GGML blocks.
It then reshapes the payload to the weight's layout. The version counter
still makes a weight edited after its forward pack afresh. The Megatron
exporter keeps packing on its own, since it can remap or slice a weight
before packing it.

**Bit-exact decoders.** One CUDA thread decodes one 8-value vector.
Every float operation is explicitly rounded (`__fmul_rn`, `__fadd_rn`,
`__fdiv_rn`) in the PyTorch decoder's order, so the compiler cannot fuse
a multiply into an add, and the output is bit-identical to the PyTorch
decoders. `dequantize_<format>` uses the unpacker for CUDA payloads and
keeps the PyTorch path otherwise.

### Testing

- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **74 passed** on RTX PRO
6000 (sm_120). That includes 25 new decoder cases: for each format, CUDA
equals the PyTorch decoder bit for bit in four dtypes, on random
payloads (including block scales that decode to inf or NaN) and real
encodings. The CUDA path also reproduces llama.cpp's values on the
conformance blocks. Dropping IQ2_XXS's parity bit in the kernel fails
all five IQ2_XXS cases.
- `tests/unit/torch/export/test_export_weight.py`: export reuses the
cached payload without calling the encoder, and repacks a weight edited
after its forward, for all five formats. The weight is 512 wide so the
blocked view really differs; with an exact-key match only, the reuse
test fails for every format. The IQ payload export test covers all five
formats rather than two.
- IQ unit tests (`test_ggml_backend.py`, `test_iq_formats.py`,
`test_convert_hf_config.py`, `test_presets.py`,
`test_export_weight.py`): **186 passed**
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **45 passed** in `nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq`: **5 passed**, 27.7–30.9
s each
- On this GPU, `test_torch_extensions.py` still shows #2515's two Q8_0
failures, which fail identically on a clean `main`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Same bindings, error messages,
exported bytes and decoded values, only faster.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. It speeds up formats added in this
unreleased cycle without changing their output.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Builds on #2615 (one CUDA encoder per IQ family), now merged. Follows
the IQ series (#2511, #2525, #2512, #2565, #2513, #2595), all merged.
The benchmark above was measured on the combined branch before the
split. After rebasing onto `main` with #2615 merged, this PR's code is
the same apart from `kVectorsPerBlock` moving into `common.cuh`'s shared
constants and #2615's validation-order fix. On the rebased branch, all
40 encoder and decoder hashes match, 74 GPU tests pass and 186 unit
tests pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hychiang-git added a commit that referenced this pull request Oct 2, 2026
### What does this PR do?

Type of change: new feature

Adds Q8_0 encoding, decoding, and the GGML fake-quant backend. Each
32-weight block stores one FP16 scale and 32 signed int8 values in 34
bytes (8.5 bits per weight).

- Generalizes dispatch to `GGML_FORMAT_REGISTRY` / `GGMLFormat` while
preserving `IQFormat` as a type alias. `IQ_FORMAT_REGISTRY` is an
IQ-only compatibility dictionary sharing the same format records, not a
mutation-propagating view.
- Retains IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, and IQ2_S registrations
alongside Q8_0.
- Uses the merged `q8_0_pack` CUDA extension when available and the
PyTorch encoder otherwise.
- Matches canonical reciprocal-then-multiply rounding, using the
unrounded FP32 scale to choose int8 values and FP16 only for serialized
scale storage.
- Adds exact-byte rounding regressions and checks that the compatibility
registry shares the same format records. Removes a redundant zero-block
buffer copy.

Checkpoint export, recipes, documentation, and the user-facing Q8_0
changelog remain in #2517.

### Base and dependencies

The target remains `main`. Current head
`71e3183a7b7eac7e5e3888ee45a206c572028af9` includes main at
`67a68f8fd4902a7b67c78a2f00f40562b081fe72`, including #2595 (IQ1_M
registration), #2615 (shared IQ CUDA encoders), and #2604 (packed-weight
cache reuse and CUDA IQ decoding). All five IQ formats, Q8_0, and the
compatibility aliases are retained.

The comparison against `main` contains only the intended 12 Q8_0 files,
with 207 added source lines excluding tests and docs. The inherited IQ
refactor and export changes are not part of this PR's diff.

#2515 has already merged and supplies the Q8_0 kernel. This PR retains
the small reciprocal-rounding correction to that kernel needed for exact
reference parity.

### Usage

```python
import torch
from modelopt.torch.quantization.ggml import quantize_q8_0, dequantize_q8_0

weight = torch.randn(2, 64, dtype=torch.bfloat16)
packed, shape = quantize_q8_0(weight)
restored = dequantize_q8_0(packed, shape)
```

The final weight dimension must be divisible by 32. Backend dispatch
also accepts `num_bits="q8_0"`, `backend="ggml"`, and `block_sizes={-1:
32}`.

### Testing

- Current head `71e3183a7`: **183 focused CPU tests passed**, covering
Q8_0, all six registered backends, IQ formats, export metadata, and
recipe presets.
- Current-head applicable pre-commit checks passed: Ruff, formatting,
mypy, CUDA formatting, license headers, security checks, merge markers,
line endings, and file size.
- Previous head `2beb70ffa`: **all six Q8_0 CUDA cases passed**,
including exact-byte reciprocal rounding and unrounded-scale
regressions; [GPU job
log](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36890153545/job/110468900810).
That GPU lane completed with 1,663 passed and 67 skipped. The overall
workflow was cancelled after a different lane was cancelled; it is not
an all-green workflow result.
- Current head `71e3183a7`: the GPU CI mirror now points to the exact PR
head. [GPU
CI](https://github.com/NVIDIA/Model-Optimizer/actions/runs/37061248628),
[example
CI](https://github.com/NVIDIA/Model-Optimizer/actions/runs/37061248555),
and [regression
CI](https://github.com/NVIDIA/Model-Optimizer/actions/runs/37061248637)
are running. Previous-head results do not validate this new head.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: yes; existing IQ names and
registrations are retained.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`?: yes; no new
dependency.
- Did you write any new necessary tests?: yes.
- Did you update `CHANGELOG.rst`?: N/A here; #2517 carries the single
Q8_0 feature entry.
- Did you get Claude approval on this PR?: pending current-head review.

### Related PRs

1. [#2515 — Q8_0 CUDA packing
kernel](#2515) — merged.
2. [#2595 — IQ1_M
registration](#2595) —
merged.
3. **#2516 — Q8_0 quantization codec and backend** — this PR.
4. [#2517 — Q8_0 checkpoint export and
recipes](#2517) — follows
this PR.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Q8_0 quantization support, including weight packing, unpacking,
and fake quantization.
* Added Q8_0 to the available GGML formats, alongside existing IQ
formats, through a shared quantization interface.
* Added support for formats with different block sizes when validating
weights.
* Q8_0 uses CUDA acceleration when available and falls back to PyTorch
when needed.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants