Skip to content

[4/5] Add the IQ2_S CUDA encoder and register the format - #2565

Merged
cjluo-nv merged 4 commits into
mainfrom
chenjiel/iq2-s-register
Sep 29, 2026
Merged

cjluo-nv merged 4 commits into
mainfrom
chenjiel/iq2-s-register

Conversation

@cjluo-nv

@cjluo-nv cjluo-nv commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Type of change: new feature

Second of two PRs adding IQ2_S (2.5625 bits per weight). #2512 landed the PyTorch codec; this PR adds its CUDA encoder and makes the format reachable:

  • the CUDA encoder, its binding and extension build wiring, plus the CUDA path in quantize_iq2_s
  • an IQFormat record and one IQ_FORMAT_REGISTRY entry, so backend dispatch, both exporters and convert_hf_config take it from there
  • the ggml package export
  • the general/ptq/iq2_s recipe, its presets, ptq.md and a CHANGELOG entry

The kernel lands with the registration so every registered format keeps a CUDA encoder.

On the mixed-precision checkpoint #2511 measured (unsloth/Qwen3.8-27B-GGUF), IQ2_S covers 9 tensors and 0.6 B parameters.

The kernel

IQ2_S's 1024-entry codebook is twice IQ2_XS's, which makes its search the most expensive in the family. The codebook and its norms take 36 KiB of shared memory, the most of any IQ kernel but inside the 48 KiB static limit, so they are declared statically like the IQ2_XS and IQ2_XXS kernels.

That cost is why the kernel matters more here than anywhere else:

torch CUDA
IQ2_S, 5632×2048 weight 0.8 M elem/s 725.7 M elem/s 907×
extrapolated to a 27B model ~9.8 hours ~37 s

Usage

python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_s

Testing

Registering the format brings it under every registry-driven test with no IQ2_S-specific test code: backend dispatch and weight caching, the num_bits guard, convert_hf_config metadata (uniform and mixed precision), all 9 Megatron export tests, and the two TensorQuantizer tests in the shared battery. The shared CUDA battery gains one row.

  • tests/unit/torch/quantization/test_ggml_backend.py, test_iq_formats.py, tests/unit/torch/export/test_convert_hf_config.py, tests/unit/recipe/test_presets.py: 134 passed
  • broader unit sweep (-k 'ggml or iq or gguf or registry' over quantization, export and recipe tests): 192 passed. The one failure, test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes, is a torchvision import error in my environment, unrelated to IQ.
  • tests/gpu/torch/quantization/test_iq_formats_cuda.py, test_iq1_s_cuda.py, test_iq2_xs_cuda.py: 42 passed on RTX PRO 6000 Blackwell (sm_120). 7 of them are IQ2_S: CUDA-vs-PyTorch encoder parity, determinism, reconstruction at scale, zero and non-finite policy, float64 input and the fallback path.
  • tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml': 36 passed (9 tests × 4 formats) in nvcr.io/nvidia/nemo:26.08
  • tests/examples/hf_ptq/test_llm_ptq.py -k iq2_s: passed. TinyLlama PTQ through unified HF export writes quant_algo: IQ2_S, block_payload_bytes: 82, and down_proj packed as (2048, 22, 82) uint8.
  • general/ptq now holds 30 recipes.
  • The shared-memory change in b7739d5d0 leaves the packed bytes identical (same hash on a 5632×2048 weight), and packing runs at 849.1 M elem/s against 825.7 before on RTX PRO 6000. The GPU battery was rerun: 42 passed.

All of the above was rerun after rebasing onto main at c2aaa44f6. That base adds a Q8_0 packer to the same GGML extension (#2515), and changes the hf_ptq example and the export code this format goes through. The packed IQ2_S bytes still hash the same. On this RTX PRO 6000 (sm_120), two of #2515's own Q8_0 tests in tests/gpu/_extensions/test_torch_extensions.py fail: test_cuda_ext_q8_0_zero_and_roundf_layout and test_cuda_ext_q8_0_dequantizes_with_small_error. They fail identically on a clean main checkout, so they are not from this PR.

Before your PR is "Ready for review"

  • Is this change backward compatible?: ✅
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: ✅ No new code sources or dependencies.
  • Did you write any new necessary tests?: ✅
  • Did you update Changelog?: ✅
  • Did you get Claude approval on this PR?: ❌ Not yet run.

Additional Information

Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) → #2512 (IQ2_S codec, merged) → this → #2513 (IQ1_M).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Added IQ2_S weight-only quantization for eligible linear layers, at 2.5625 bits per weight.
    • Added a PTQ recipe that requires no calibration data. Weights must meet the existing 256-value block-size constraint.
    • Added CUDA-accelerated packing for CUDA weights, with a Python fallback when the CUDA extension is unavailable.
  • Documentation
    • Updated the PTQ recipe catalog and IQ-format size tradeoffs.

@cjluo-nv
cjluo-nv requested review from a team as code owners September 28, 2026 16:46
@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: e88ef486-06ba-4414-82b2-1890bbf4a00d

📥 Commits

Reviewing files that changed from the base of the PR and between 86b1021ef69b1897da1fa36623faf7e3bc242196 and b7739d5.

📒 Files selected for processing (9)
  • CHANGELOG.rst
  • modelopt/torch/export/quant_format.py
  • modelopt/torch/kernels/quantization/ggml/ggml.cpp
  • modelopt/torch/quantization/extensions.py
  • modelopt/torch/quantization/ggml/__init__.py
  • modelopt/torch/quantization/ggml/iq2_s.py
  • modelopt/torch/quantization/ggml/registry.py
  • tests/gpu/torch/quantization/test_iq_formats_cuda.py
  • tests/unit/recipe/test_presets.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Walkthrough

Walkthrough

Adds IQ2_S weight-only quantization with a CUDA packer and a general/ptq recipe. IQ2_S packs blocks of 256 values at 2.5625 bits per weight. Tests and documentation now include IQ2_S.

Changes

IQ2_S quantization

Layer / File(s) Summary
CUDA block packer
modelopt/torch/kernels/quantization/ggml/common.cuh, modelopt/torch/kernels/quantization/ggml/ggml.cpp, modelopt/torch/kernels/quantization/ggml/iq2_s.cu, modelopt/torch/quantization/extensions.py, tests/gpu/torch/quantization/test_iq_formats_cuda.py
Adds the IQ2_S payload layout and CUDA encoder. The extension wrapper validates tensors and scales before packing. The CUDA test matrix includes IQ2_S.
Quantization format integration
modelopt/torch/export/quant_format.py, modelopt/torch/quantization/ggml/iq2_s.py, modelopt/torch/quantization/ggml/__init__.py, modelopt/torch/quantization/ggml/registry.py
Adds the IQ2_S format constant and registry entry. CUDA weights use the GGML extension when available; other cases use the existing Python encoder.
PTQ recipe and validation
modelopt_recipes/configs/numerics/iq2_s.yaml, modelopt_recipes/configs/ptq/presets/model/iq2_s.yaml, modelopt_recipes/general/ptq/iq2_s.yaml, modelopt_recipes/ptq.md, tests/unit/recipe/test_presets.py, tests/examples/hf_ptq/test_llm_ptq.py, CHANGELOG.rst
Adds an IQ2_S weight-only PTQ recipe for eligible linear layers, without calibration. Updates the recipe catalog and IQ format size documentation. Adds recipe contract and example test coverage, and records the release notes.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant quantize_iq2_s
  participant iq2_s_pack
  participant iq2_s_pack_cuda
  quantize_iq2_s->>iq2_s_pack: pass input, grid, and predicted scales
  iq2_s_pack->>iq2_s_pack_cuda: pass contiguous tensors
  iq2_s_pack_cuda-->>iq2_s_pack: return packed blocks
Loading

Suggested reviewers: hychiang-git

Merge Risk: ⚪ Minimal · up to b7739

The IQ2_S encoder, format integration, and PTQ recipe are consistent across the reviewed boundaries. No actionable merge risk was established.

🚥 Pre-merge checks | ✅ 5 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 61.11% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 18 functions across 13 files. (1 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Security Anti-Patterns ✅ Passed PASS. The authoritative diff adds no listed security anti-pattern. Added Python lines contain no unsafe torch.load or numpy.load settings, hardcoded trust_remote_code=True, eval/exec on external input…
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely summarizes the main changes: adding the IQ2_S CUDA encoder and registering the format. The “[4/5]” prefix does not obscure the change.
Full details: Docstring Coverage

Explanation

Docstring coverage is 61.11% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 18 functions across 13 files. (1 skipped: 1 unsupported.)

✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-09-29 20:36 UTC

@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 78.81%. Comparing base (834c90d) to head (31d0637).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2565      +/-   ##
==========================================
+ Coverage   69.40%   78.81%   +9.41%     
==========================================
  Files         610      610              
  Lines       68087    68101      +14     
==========================================
+ Hits        47256    53675    +6419     
+ Misses      20831    14426    -6405     
Flag Coverage Δ
examples-diffusers 21.36% <57.14%> (-0.04%) ⬇️
examples-gpt-oss 13.58% <57.14%> (+0.07%) ⬆️
examples-hf_ptq 23.22% <100.00%> (+0.09%) ⬆️
examples-llm_distill 13.65% <57.14%> (+0.07%) ⬆️
examples-llm_eval 17.60% <57.14%> (+0.06%) ⬆️
examples-llm_qat 17.77% <57.14%> (+0.05%) ⬆️
examples-llm_sparsity 16.04% <57.14%> (+0.07%) ⬆️
examples-megatron_bridge 26.60% <57.14%> (-0.08%) ⬇️
examples-specdec_bench 13.35% <57.14%> (+0.07%) ⬆️
examples-speculative_decoding 17.97% <57.14%> (+0.01%) ⬆️
examples-torch_onnx 21.92% <57.14%> (+0.05%) ⬆️
examples-torch_trt 15.39% <57.14%> (+0.07%) ⬆️
examples-vllm_serve 13.83% <57.14%> (+0.07%) ⬆️
gpu 58.43% <100.00%> (+36.94%) ⬆️
regression 15.28% <57.14%> (+0.07%) ⬆️
unit 59.10% <64.28%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@cjluo-nv
cjluo-nv force-pushed the chenjiel/iq2-s-register branch from 1e6d33f to 6b1c38c Compare September 28, 2026 17:09
@cjluo-nv cjluo-nv changed the title [4/5] Register the IQ2_S format and add its recipe [4/5] Add the IQ2_S CUDA encoder and register the format Sep 28, 2026
cjluo-nv added a commit that referenced this pull request Sep 28, 2026
### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ2_S**, the widest of the GGML IQ formats at
one and two bits (2.5625 bits per weight). This one lands the **PyTorch
codec**: the encoder, the decoder and the 1024-entry codebook. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2565 adds the CUDA encoder,
registers the format and adds its recipe.

### What's distinctive about it

**IQ2_S is the one format llama.cpp's own tooling gives no head start
on**, so the search is written against the GGML layout directly.

The interesting difference from IQ2_XS and IQ2_XXS is sign handling.
IQ2_S stores a **full 8-bit sign mask** per group rather than a 7-bit
parity-coded index. The encoder therefore takes the input signs as they
are instead of flipping the weakest element to fix parity, and the
search compares magnitudes directly, which is simpler than its siblings.

### Why the codec lands before the kernel

The CUDA encoder's tests use this codec as their reference. They compare
against the PyTorch encoder byte for byte and draw the grid and scale
predictor from it. So the kernel cannot be tested before the codec
exists, and it follows in #2565 together with the registration. Every
registered format therefore keeps a CUDA encoder.

### Test changes that make the split possible

A codec can now land before it is registered, so two test contracts in
`test_iq_formats.py` are stated precisely:

- The two tests that go through `TensorQuantizer` (pass-through
gradient, error falls with bit width) iterate `IQ_FORMAT_REGISTRY`.
Every other battery test calls the codec directly and covers IQ2_S here.
- The coverage check now asserts `set(IQ_FORMAT_REGISTRY) <=
set(FORMATS)` instead of equality. That is what its docstring already
said: a registered format must be listed, or it escapes the contract.
- `test_registry_lists_every_exported_encoder` is unchanged, and it is
why this PR leaves the package exports alone: an exported encoder must
be registered.

The error-by-bit-width failure message also labels errors by the order
they were measured in; it previously zipped them with alphabetical
names.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ2_S: 9 tensors, 2,355,200 blocks → 0 mismatched, max|diff| 0.0
```

The new codebook matches the `ggml-common.h` table entry for entry.
Blocks from `unsloth/Qwen3.8-27B-GGUF` ship as conformance vectors, so
CI keeps checking bytes we did not produce.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **121 passed**, 14 of them IQ2_S
codec cases, including the llama.cpp conformance check
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **35 passed**, unchanged by
this PR

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ The new
codebook is a GGML table, carried in `codebooks.py` with the source
revision recorded. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2565
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) →
**this** → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added GGML-compatible IQ2_S quantization and dequantization support,
including access to its magnitude grid.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cjluo-nv
cjluo-nv force-pushed the chenjiel/iq2-s-register branch from 6b1c38c to 70d8666 Compare September 28, 2026 21:07

@meenchen meenchen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bot review (gpt-6-astra) — DM the bot to share feedback.

Nudge: the integration and shared tests look sound, but the shared-memory rationale is incorrect and the GGML source attribution needs human confirmation.

Needs action:

  • Correct the shared-memory claim in iq2_s.cu and the PR body: the grid plus norms uses 36 KiB, and scratch brings the total below 48 KiB. Use static storage or explain the actual dynamic-storage benefit.
  • Confirm that iq2_s.cu references GGML only as a format specification, with no copied implementation requiring additional third-party notices. Its NVIDIA header matches LICENSE_HEADER.

No action needed:

  • The design extends existing IQFormat dispatch and CUDA helpers; the reported performance comparison justifies accelerating the existing PyTorch codec rather than adding another subsystem.
  • Test edits extend parametrization without weakening assertions. Shared coverage includes encoder parity, determinism, reconstruction, fallback, registry wiring, caching, and export metadata. Tests were inspected, not executed.

@cjluo-nv

Copy link
Copy Markdown
Collaborator Author

On the shared-memory item in the review above: right, the grid plus norms is 36 KiB and fits the 48 KiB static limit. 86b1021ef declares them statically, as the IQ2_XS and IQ2_XXS kernels do, and drops the per-call cudaFuncSetAttribute. The packed bytes are unchanged (same hash on a 5632×2048 weight), and packing is about 3% faster. The PR description is corrected.

@cjluo-nv

Copy link
Copy Markdown
Collaborator Author

On the GGML provenance item: iq2_s.cu uses GGML only as the format specification. From GGML it takes the 82-byte block_iq2_s layout (field offsets, index and sign bit packing, and the d·(2·ls+1)/8 local scale), cited by revision at the top of the file.

The search is not llama.cpp's. quantize_row_iq2_s_impl tries 19 trial scales per 16-value group, maps rounded levels through kmap_q2xs and falls back to iq2_find_best_neighbour, with importance weights. This kernel searches all 1024 entries exhaustively at a fixed predicted block scale, and shares none of those routines or identifiers.

The codebook values are not in the kernel either: they are passed in from codebooks.py, which #2512 added with the upstream revision recorded. No additional third-party notice is needed for this file.

@meenchen meenchen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bot review (gpt-6-astra) — DM the bot to share feedback.

LGTM: prior technical concerns are resolved; human sign-off remains for GGML provenance and the justified test extensions.

Needs action:

  • Sign off on the GGML format-specification attribution in iq2_s.cu and the expanded CUDA, recipe, and PTQ test parametrization.

No action needed:

  • ✔️ Resolved since the last review: static shared storage and corrected 36 KiB accounting; the author specifically explained the independent search implementation and specification-only GGML reference.
  • The NVIDIA header matches LICENSE_HEADER. No codebook or upstream search implementation is added in this diff.
  • Design remains sound: this extends existing IQFormat dispatch and CUDA helpers, retaining the PyTorch fallback. The reported performance comparison justifies acceleration.
  • Test edits add IQ2_S coverage without weakening assertions. Shared tests cover parity, determinism, reconstruction, zero/non-finite inputs, float64, fallback, caching, and export metadata. Tests were inspected, not executed.

@meenchen meenchen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Didn't look at the cuda part closer, but others LGTM

Comment thread modelopt_recipes/ptq.md
Comment on lines 62 to +64
| `iq2_xxs` | IQ2_XXS W2A16 (2.06 bpw), eligible linears | none | none (no calibration) |
| `iq2_xs` | IQ2_XS W2A16 (2.31 bpw), eligible linears | none | none (no calibration) |
| `iq2_s` | IQ2_S W2A16 (2.56 bpw), eligible linears | none | none (no calibration) |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we plan to have W2A4 for IQ formats in the future?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes

cjluo-nv and others added 2 commits September 29, 2026 06:40
Second of two changes adding IQ2_S. The previous change landed the PyTorch
codec; this one adds its CUDA encoder and makes the format reachable.

The 1024-entry codebook is twice IQ2_XS's, which makes IQ2_S the most
expensive search in the family and pushes the grid past the static shared
memory limit, so the kernel keeps the codebook and its norms in dynamic shared
memory. On a 5632x2048 weight that is 725.7 M elem/s against the torch
search's 0.8 -- without the kernel, a 27B model would take about ten hours to
pack. The kernel is checked byte for byte against the PyTorch encoder.

IQ2_S gets an IQFormat record and one IQ_FORMAT_REGISTRY entry, so backend
dispatch, both exporters and convert_hf_config take it from there. The ggml
package exports it, and a general/ptq/iq2_s recipe uses it. The kernel lands
with the registration so every registered format keeps a CUDA encoder.

IQ2_S covers 9 tensors and 0.6B parameters of the mixed-precision checkpoint
IQ2_XXS measured. Registering it brings it under every registry-driven test
with no IQ2_S-specific test code: backend dispatch and weight caching, the
num_bits guard, convert_hf_config metadata and Megatron export. The hf_ptq
example test gains an IQ2_S case, which the CUDA encoder keeps within that
test's time budget.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
The kernel put the 1024-entry codebook and its norms in dynamic shared memory
on the grounds that they exceed the static limit. They do not: 1024 x 8
floats plus 1024 norms is 36 KiB, and the kernel's other shared arrays add
well under 1 KiB, inside the 48 KiB static limit. Declare them statically, as
the IQ2_XS and IQ2_XXS kernels do, and drop the launch-time size calculation
and the cudaFuncSetAttribute call on every pack.

The packed bytes are unchanged: a 5632x2048 weight hashes the same before and
after. Packing is about 3% faster (825.7 to 849.1 M elem/s on an RTX PRO 6000)
without the per-call attribute set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
@cjluo-nv
cjluo-nv force-pushed the chenjiel/iq2-s-register branch from 86b1021 to b7739d5 Compare September 29, 2026 06:46
@cjluo-nv
cjluo-nv enabled auto-merge (squash) September 29, 2026 06:49
@cjluo-nv
cjluo-nv merged commit 3091b8f into main Sep 29, 2026
55 checks passed
@cjluo-nv
cjluo-nv deleted the chenjiel/iq2-s-register branch September 29, 2026 20:35
cjluo-nv added a commit that referenced this pull request Sep 30, 2026
### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ1_M** at 1.75 bits per weight, just above
IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2595 adds the CUDA encoder,
registers the format and adds its recipe. With both, ModelOpt supports
all five GGML IQ formats at one and two bits.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B
parameters**. With all five formats we can read 89.0% of that file; the
rest is k-quants and F32.

### What's distinctive about it

**IQ1_M is the most irregular layout of the five.** There is no leading
block scale field at all. The FP16 super-block scale is reassembled from
the top nibble of each of four scale words:

```c
scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000);
```

It is also finer grained than IQ1_S: a local scale per **two** groups
rather than four, and a delta shift chosen **per group** rather than per
sub-block. That is where its extra 0.1875 bits go.

### Shared with IQ1_S rather than copied

IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same
±1/8 delta, every (shift, local scale) choice for every 8-value vector.
It differs only in how it selects among those choices afterwards. So the
search moves out of IQ1_S's encoder into `_search_shifted_grid`, which
both call, and `iq1_m.py` keeps only its selection and packing.
**IQ1_S's encoded bytes are unchanged**, checked by hashing its output
before and after on a fixed input.

### A scale-anchor correction

IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a
block's peak-to-RMS** rather than being flat, and clamps higher. It uses
`clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat
`0.61`. Measured over 15 Qwen3.8-27B MLP weights:

| | flat 0.61 | correct anchor | |
|---|---|---|---|
| relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** |

It is consistent on every tensor, with no outliers. The anchor changes
quality without touching layout, so neither round-trip nor conformance
tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms`
now pins it, for all five formats; see Testing.

### Family parity

Two surface asymmetries close here, so the five are uniform. `IQ1_S` now
exposes `_predict_iq1_s_scales` like the other four, instead of
computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the
IQ1_S table it shares.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0
```

This mattered: **my first IQ1_M decoder had a real bug.** A
`repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where
llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder
still passed, because the encoder made the matching mistake. Only
comparison against bytes we did not produce caught it. Blocks from that
checkpoint ship as conformance vectors, and mutation testing confirms
they catch a mis-set scale nibble.

The decoder unpacks every field in one vectorized pass, since fake quant
decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3).

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M
codec cases, including the llama.cpp conformance check
- `test_scale_anchor_follows_peak_to_rms` pins every format's scale
anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4,
8 and 16, reaching both clamps and two points on each slope, and
compares them against anchors written out in the test. Mutations each
fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61,
moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S
or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's
CUDA-vs-PyTorch parity still holds after its encoder refactor.
- IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash
identically to the pre-split version of this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds
no codebook; it reuses the IQ1_S table already carried in
`codebooks.py`. The new conformance vectors come from
`unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model
`Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration), all merged →
**this** → #2595 (IQ1_M CUDA encoder and registration).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M quantization and dequantization for compact,
GGML-compatible blocks of 256 values.
* Added access to the IQ1_M grid and configurable chunk sizes for
processing data.

* **Bug Fixes**
* Improved IQ1_S scale prediction and grid-search organization while
preserving its encoding behavior.

* **Tests**
* Added IQ1_M conformance data and included the format in shared
IQ-format test coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
cjluo-nv added a commit that referenced this pull request Oct 1, 2026
### What does this PR do?

Type of change: new feature

**Second of two PRs adding IQ1_M** (1.75 bits per weight). #2513 landed
the PyTorch codec; this PR adds its **CUDA encoder** and makes the
format reachable. With it, ModelOpt supports all five GGML IQ formats at
one and two bits.

- the CUDA encoder, its binding and extension build wiring, plus the
CUDA path in `quantize_iq1_m`
- an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so
backend dispatch, both exporters and `convert_hf_config` take it from
there
- the `ggml` package export
- the `general/ptq/iq1_m` recipe, its presets, `ptq.md` and a CHANGELOG
entry

The kernel lands with the registration so every registered format keeps
a CUDA encoder.

### The kernel

In the kernel the delta shift is free per group, so it sits above the
entry index in the sort key: a tie still prefers the lower shift and
then the lower entry, as the reference encoder does. The 2048-entry grid
IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory
limit, so both kernels read it from global memory and rely on the cache.

| 5632×2048 weight | torch | CUDA | |
|---|---|---|---|
| IQ1_M encode | 5.6 M elem/s | **318 M elem/s** | **57×** |

### Shared with IQ1_S rather than copied

The two IQ1 kernels load each vector, score it against a grid entry and
apply the ±1/8 shift the same way. So those three steps move into
`common.cuh` as `load_vector`, `grid_terms` and `shifted_error`, and
IQ1_S uses them too. **IQ1_S's packed bytes are unchanged**: its CUDA
output on a 5632×2048 weight hashes the same before and after, and so
does IQ1_M's, compared against the pre-split version of this change.
IQ1_S encodes at the same speed (306 M elem/s).

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m
```

### Testing

Registering the format brings it under every registry-driven test with
no IQ1_M-specific test code: backend dispatch, weight caching, the
`num_bits` guard, `convert_hf_config` metadata, Megatron export and the
`TensorQuantizer` tests in the shared battery. The shared CUDA battery
gains one row.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **166 passed**
- broader unit sweep (`-k 'ggml or iq or gguf or registry'` over
quantization, export and recipe tests): **221 passed**. The one failure,
`test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`,
is a `torchvision` import error in my environment, unrelated to IQ.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** on RTX PRO
6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch
encoder parity
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **45 passed** (9 tests × 5 formats) in
`nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m`: **passed**
- reconstruction error falls monotonically across all five formats,
pinned by a test
- `general/ptq` now holds 31 recipes; `ptq.md` is updated.

Rebased onto `main` after #2513 merged. The resulting tree is identical
to the one the runs above tested, and the unit set was rerun on it: 166
passed.

On this GPU, two of #2515's Q8_0 tests in
`tests/gpu/_extensions/test_torch_extensions.py` fail:
`test_cuda_ext_q8_0_zero_and_roundf_layout` and
`test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically
on a clean `main` checkout, so they are not from this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M
codec), all merged → **this**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M weight-only quantization at 1.75 bits per weight, with
CUDA acceleration and a 256-value block size.
* Added an IQ1_M post-training quantization recipe for eligible linear
layers; calibration data is not required.
* Added IQ1_M to the supported GGML-compatible formats and recipe
listings.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
cjluo-nv added a commit that referenced this pull request Oct 2, 2026
### What does this PR do?

Type of change: performance

Two fixes that make fake-quantized GGML IQ models fast. Both were found
by instrumenting the hf_ptq example (TinyLlama, 154 IQ weights,
100-token preview) and counting every encode and decode per weight.

1. **Each weight is packed once.** Fake quant packs a weight on its
first forward and caches the payload, but export then ran the search
again on the same unchanged weight: 154 more encodes for 154 weights, in
all five formats. `IQFormat.pack` now reuses the cached payload when it
was packed from the weight being exported, so the checkpoint also holds
exactly the bytes the evaluated model decoded.
2. **Decoding runs on CUDA.** All five formats had CUDA encoders but
decoded with PyTorch ops. Fake quant decodes every weight on every
forward, so once packing was fast and cached, decoding was **55–66% of
the run**, about 3 ms per weight per forward; on a 27B model that is
about 10 s per forward. Each format's `Format` struct from #2615 gains a
bit-exact `decode()` beside its `store()`.

### Results

The instrumented hf_ptq example on the same GPU, with the extension
already built:

| format | total wall before → after | encodes at export | decode time
before → after |
|---|---|---|---|
| IQ1_S | 55.6 → **21.9 s** | 154 → **0** | 32.0 → **1.4 s** |
| IQ1_M | 69.8 → **21.0 s** | 154 → **0** | 46.3 → **1.3 s** |
| IQ2_XXS | 61.6 → **19.0 s** | 154 → **0** | 41.9 → **1.3 s** |
| IQ2_XS | 63.8 → **19.1 s** | 154 → **0** | 44.3 → **1.3 s** |
| IQ2_S | 55.6 → **19.9 s** | 154 → **0** | 35.8 → **1.3 s** |

Fake quant still packs each weight exactly once and decodes it 100
times, once per preview token. IQ1_M was the slowest format before
because its PyTorch decoder did the most work; it now matches IQ1_S.
Decoding one 5632×2048 weight goes from 3.3–5.2 ms to **0.06–0.14 ms**.

### How

**Reusing the payload.** TensorQuantizer hands fake quant the weight
reshaped into 256-value blocks, so the cache is keyed on that view,
while export holds the weight itself. An exact-key match therefore never
hit in a real model, even though a 256-wide unit test passed. `pack`
also accepts another contiguous view of the same storage, version and
length: the same values in the same order, hence the same GGML blocks.
It then reshapes the payload to the weight's layout. The version counter
still makes a weight edited after its forward pack afresh. The Megatron
exporter keeps packing on its own, since it can remap or slice a weight
before packing it.

**Bit-exact decoders.** One CUDA thread decodes one 8-value vector.
Every float operation is explicitly rounded (`__fmul_rn`, `__fadd_rn`,
`__fdiv_rn`) in the PyTorch decoder's order, so the compiler cannot fuse
a multiply into an add, and the output is bit-identical to the PyTorch
decoders. `dequantize_<format>` uses the unpacker for CUDA payloads and
keeps the PyTorch path otherwise.

### Testing

- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **74 passed** on RTX PRO
6000 (sm_120). That includes 25 new decoder cases: for each format, CUDA
equals the PyTorch decoder bit for bit in four dtypes, on random
payloads (including block scales that decode to inf or NaN) and real
encodings. The CUDA path also reproduces llama.cpp's values on the
conformance blocks. Dropping IQ2_XXS's parity bit in the kernel fails
all five IQ2_XXS cases.
- `tests/unit/torch/export/test_export_weight.py`: export reuses the
cached payload without calling the encoder, and repacks a weight edited
after its forward, for all five formats. The weight is 512 wide so the
blocked view really differs; with an exact-key match only, the reuse
test fails for every format. The IQ payload export test covers all five
formats rather than two.
- IQ unit tests (`test_ggml_backend.py`, `test_iq_formats.py`,
`test_convert_hf_config.py`, `test_presets.py`,
`test_export_weight.py`): **186 passed**
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **45 passed** in `nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq`: **5 passed**, 27.7–30.9
s each
- On this GPU, `test_torch_extensions.py` still shows #2515's two Q8_0
failures, which fail identically on a clean `main`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Same bindings, error messages,
exported bytes and decoded values, only faster.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. It speeds up formats added in this
unreleased cycle without changing their output.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Builds on #2615 (one CUDA encoder per IQ family), now merged. Follows
the IQ series (#2511, #2525, #2512, #2565, #2513, #2595), all merged.
The benchmark above was measured on the combined branch before the
split. After rebasing onto `main` with #2615 merged, this PR's code is
the same apart from `kVectorsPerBlock` moving into `common.cuh`'s shared
constants and #2615's validation-order fix. On the rebased branch, all
40 encoder and decoder hashes match, 74 GPU tests pass and 186 unit
tests pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants