Skip to content

[OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [4/5] - #2579

Merged
shengliangxu merged 0 commit into
shengliangx/move-modeling_3from
shengliangx/move-modeling_4
Oct 6, 2026
Merged

shengliangxu merged 0 commit into
shengliangx/move-modeling_3from
shengliangx/move-modeling_4

Conversation

@shengliangxu

@shengliangxu shengliangxu commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Type of change: Refactor

Summary of the series

Model-specific PTQ modeling moves out of modelopt/torch/quantization/plugins/huggingface.py into the per-model packages under modelopt/torch/models/<model_type>/modeling_ptq.py, next to each model's specs.py. By the end of the series huggingface.py shrinks from ~2145 to ~1490 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention, FP8Linear, CompressedLinear, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had no ModelSpec get one.

  • Class and function bodies move verbatim. The only rewrites: on-the-fly callbacks are added to CUSTOM_MODEL_PLUGINS by their own module, and is_homogeneous_hf_model imports is_nemotron_h_model lazily.
  • huggingface.py imports every modeling_ptq from an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first.
  • Each package's __init__.py keeps importing only specs, so import modelopt.torch.models does not pull in quantization or transformers.
  • No public API or quantization behavior changes; only private (_-prefixed) names change module.

Merge in order. #2576 targets main; each later PR targets the one before it.

  1. [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [1/4] #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and step3p7
  2. [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [2/4] #2577 — Add the modeling_ptq.py convention; move Nemotron-H, Falcon and Llama4
  3. [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [3/4] #2578 — Move GPT-OSS and Qwen3-VL-MoE
  4. [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [4/5] #2579 — Move DBRX ← this PR
  5. [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [4/4] #2580 — Move the Step family (step3p5, shared by step3p7)

This PR [4/5]

Moves _QuantDbrxExperts, _QuantDbrxExpertGLU, _QuantDbrxFFN, their registrations and register_dbrx_moe_on_the_fly into dbrx/modeling_ptq.py. The DBRX comment in dbrx/specs.py and the customized-model quantization guide point at the new file.

Usage

No API change. To add PTQ support for a new model, put its wrapper in modelopt/torch/models/<model_type>/modeling_ptq.py and add <model_type> to the import loop at the end of quantization/plugins/huggingface.py.

Testing

Each branch of the stack checked out and tested with pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe on CPU (torch 2.11, transformers 5.14), on main at e3c903c (the #2569 squash): 2155 passed, 7 skipped, plus 1 failure in test_autoquant_shapley.py::test_data_parallel_aumann_shapley, a multi-process test that is flaky on main itself (it failed 2 of 4 runs on main with none of these changes). The stack has since been rebased onto main at 44bbd47; the commits in between do not touch code under this change (the last one only edits AGENTS.md). GPU tests and the transformers 4.57 matrix are left to CI.

Before your PR is "Ready for review"

  • Is this change backward compatible?: ✅ Only private (_-prefixed) names change module; public API and quantization behavior are unchanged.
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: N/A
  • Did you write any new necessary tests?: N/A (verbatim move; covered by existing tests)
  • Did you update Changelog?: N/A (internal refactor)
  • Did you get Claude approval on this PR?: ❌

Additional Information

Part of a 5-PR stack; see the series list above.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Added post-training quantization support for DBRX models, including models with dynamically loaded expert layers. Support covers different DBRX routing configurations, including legacy and newer settings.
  • Documentation
    • Updated the DBRX quantization guide to reference the current location of the support.

@copy-pr-bot

copy-pr-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: 5bce6a82-0eb9-4110-964c-f3ff90e33dfc
📥 Commits

Reviewing files that changed from the base of the PR and between aae8551 and 9a18923.

📒 Files selected for processing (2)
  • modelopt/torch/models/dbrx/specs.py
  • modelopt/torch/quantization/plugins/huggingface.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • modelopt/torch/models/dbrx/specs.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.


📝 Walkthrough

Walkthrough

DBRX PTQ wrappers and dynamic expert registration now reside in a DBRX-specific module. The Hugging Face plugin imports that module through its model-type loop. The guide and specs comment point to the new implementation path.

Changes

DBRX PTQ

Layer / File(s) Summary
DBRX quantized wrappers
modelopt/torch/models/dbrx/modeling_ptq.py, modelopt/torch/quantization/plugins/huggingface.py
The DBRX expert routing, expert GLU, and FFN wrappers are defined in modeling_ptq.py. Their former definitions are removed from the Hugging Face plugin.
DBRX registration and references
modelopt/torch/models/dbrx/modeling_ptq.py, modelopt/torch/quantization/plugins/huggingface.py, modelopt/torch/models/dbrx/specs.py, docs/source/guides/_customized_model_quantization.rst
The DBRX module registers Transformers wrappers and dynamically loaded expert GLU classes. The Hugging Face plugin imports the module through its model-type loop. The guide and specs comment point to the new implementation path.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Refactor

Merge Risk: ⚪ Minimal · up to 9a189

The DBRX PTQ relocation has no identified behavior change that needs fixing before merge. Normal validation can proceed.

🚥 Pre-merge checks | ✅ 5 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 46.67% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 15 functions across 3 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the main change: moving model-specific PTQ code into modelopt/torch/models. The [4/5] suffix adds minor noise but does not make the title unclear.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Security Anti-Patterns ✅ Passed No listed security anti-pattern was introduced. The only Python addition is modelopt/torch/models/dbrx/modeling_ptq.py; its added code contains no torch.load(..., weights_only=False), `numpy.load(…
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 88.00000% with 9 lines in your changes missing coverage. Please review.
✅ Project coverage is 78.42%. Comparing base (44fdb8d) to head (9a18923).

Files with missing lines Patch % Lines
modelopt/torch/models/dbrx/modeling_ptq.py 88.00% 9 Missing ⚠️
Additional details and impacted files
@@                       Coverage Diff                       @@
##           shengliangx/move-modeling_3    #2579      +/-   ##
===============================================================
+ Coverage                        76.16%   78.42%   +2.25%     
===============================================================
  Files                              635      636       +1     
  Lines                            69702    69709       +7     
===============================================================
+ Hits                             53090    54670    +1580     
+ Misses                           16612    15039    -1573     
Flag Coverage Δ
examples-diffusers 21.13% <40.00%> (+<0.01%) ⬆️
examples-hf_ptq 22.94% <40.00%> (+<0.01%) ⬆️
examples-llm_distill 13.58% <38.66%> (+<0.01%) ⬆️
examples-llm_eval 17.48% <40.00%> (+<0.01%) ⬆️
examples-llm_qat 17.62% <40.00%> (+<0.01%) ⬆️
examples-llm_sparsity 15.92% <38.66%> (+<0.01%) ⬆️
examples-megatron_bridge 26.84% <40.00%> (+<0.01%) ⬆️
examples-specdec_bench 13.29% <38.66%> (+<0.01%) ⬆️
examples-speculative_decoding 17.82% <40.00%> (+<0.01%) ⬆️
examples-torch_onnx 21.68% <40.00%> (+<0.01%) ⬆️
examples-torch_trt 15.30% <40.00%> (+<0.01%) ⬆️
examples-vllm_serve 13.76% <38.66%> (+<0.01%) ⬆️
gpu 58.39% <40.00%> (+8.77%) ⬆️
regression 15.17% <38.66%> (+<0.01%) ⬆️
unit 58.67% <88.00%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_4 branch from a989c02 to c21bc33 Compare October 2, 2026 01:12
@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

QR code for preview link

🚀 View preview at
https://NVIDIA.github.io/Model-Optimizer/pr-preview/pr-2579/

Built to branch gh-pages at 2026-10-06 06:23 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_3 branch from 82fe8d0 to cc9c1c2 Compare October 2, 2026 17:13
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_4 branch from c21bc33 to 28d7ec6 Compare October 2, 2026 17:13
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_3 branch from cc9c1c2 to a1f35f4 Compare October 2, 2026 17:17
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_4 branch from 28d7ec6 to 820d490 Compare October 2, 2026 17:17
@shengliangxu
shengliangxu marked this pull request as ready for review October 2, 2026 17:19
@shengliangxu
shengliangxu requested review from a team as code owners October 2, 2026 17:19
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_3 branch from a1f35f4 to 561aa83 Compare October 2, 2026 17:37
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_4 branch 3 times, most recently from aae8551 to 9a18923 Compare October 6, 2026 06:17
@shengliangxu
shengliangxu removed this pull request from stack #2631 October 6, 2026 09:15
@shengliangxu
shengliangxu merged commit f29052e into shengliangx/move-modeling_3 Oct 6, 2026
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_3 branch from 44fdb8d to f29052e Compare October 6, 2026 09:15
@shengliangxu
shengliangxu deleted the shengliangx/move-modeling_4 branch October 6, 2026 09:15
@shengliangxu
shengliangxu force-pushed the shengliangx/move-modeling_4 branch from 9a18923 to f29052e Compare October 6, 2026 09:15
@shengliangxu

Copy link
Copy Markdown
Collaborator Author

This slice is dropped from the series: DBRX support was removed in #2649, so there is no DBRX PTQ modeling left to move. Its branch now equals shengliangx/move-modeling_3 (no commits of its own), which is why GitHub shows this PR as merged — nothing was added to any branch. The stack is now #2576 → #2577 → #2578 → #2580, renumbered [1/4]–[4/4].

shengliangxu added a commit that referenced this pull request Oct 6, 2026
)

### What does this PR do?

Type of change: Refactor

### Summary of the series

Model-specific PTQ modeling moves out of
`modelopt/torch/quantization/plugins/huggingface.py` into the per-model
packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`,
next to each model's `specs.py`. By the end of the series
`huggingface.py` shrinks from ~2025 to ~1525 lines and keeps only the
code shared across models: sequential/fused MoE auto-detection,
attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the
transposed-quantization helpers that gpt_oss and llama4 share. Model
types that gain PTQ modeling but had no `ModelSpec` get one.

- Class and function bodies move verbatim. The only rewrites: on-the-fly
callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and
`is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily.
- `huggingface.py` imports every `modeling_ptq` from an explicit list.
The list sits after the generic wrappers these modules build on and
before the homogeneous decoder discoverer is registered, so Nemotron-H's
more specific discoverer still matches first.
- Each package's `__init__.py` keeps importing only `specs`, so `import
modelopt.torch.models` does not pull in quantization or transformers.
- No public API or quantization behavior changes; only private
(`_`-prefixed) names change module.

Merge in order. #2576 targets `main`; each later PR targets the one
before it.

1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and
step3p7 ← **this PR**
2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon
and Llama4
3. #2578 — Move GPT-OSS and Qwen3-VL-MoE
4. #2580 — Move the Step family (step3p5, shared by step3p7)

The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there
was nothing left to move.

### This PR [1/4]

Registers a `ModelSpec` for each model type that gains `modeling_ptq.py`
later in the series but has no spec yet. Each spec records only facts
checked against transformers 4.57 and 5.14:

- `llama4` and `qwen3_vl_moe`: fused MoE layout
(`gate_up_proj`/`down_proj`), no gate/up pair, grouped export off. This
is data only: both blocks were already detected as MoE (structurally and
by name), and their expert containers are exported without name lookups.
- `falcon`: dense, no sections.
- `step3p5` / `step3p7`: `modeling_source="remote_code"` and
intentionally no `MoESpec`. Step's expert projections sit directly on
the MoE MLP with no `experts` container. Declaring the block would make
`is_moe` claim it and send AWQ export into `get_experts_list`, which
does not support that layout. `step3p7` gets its own package, following
the `gemma4_text` / `gemma4` precedent.

The exhaustive spec tables in
`tests/unit/torch/models/test_model_specs.py` (`EXPECTED_MOE_LAYOUTS`,
`root_class_names`) gain the two MoE rows.

### Usage

No API change. The new specs are read through the existing registry,
e.g. `get_spec("llama4").moe_spec`.

### Testing

Each branch of the stack checked out and tested with `pytest
tests/unit/torch/models tests/unit/torch/quantization
tests/unit/torch/export tests/unit/recipe` on CPU (torch 2.11,
transformers 5.14), on `main` at 9f902ae (#2649): 2248 passed, 7
skipped. The new block names resolve in `test_specs_vs_transformers.py`,
and the step3p5/step3p7 remote-code absence checks pass. Class names
were also checked against the transformers 4.57 wheel. GPU tests and the
transformers 4.57 matrix are left to CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Only private (`_`-prefixed)
names change module; public API and quantization behavior are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ Spec tables extended
(`EXPECTED_MOE_LAYOUTS`, `root_class_names`).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal refactor)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Part of a 5-PR stack; see the series list above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added model support for Falcon, Llama 4, Qwen3-VL-MoE, Step-3.5, and
Step-3.7.
* Llama 4 and Qwen3-VL-MoE support includes recognition of their fused
expert layers.
* Falcon, Llama 4, and Qwen3-VL-MoE require Transformers 4.57 or later.
Step-3.5 and Step-3.7 use remote-code modeling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
shengliangxu added a commit that referenced this pull request Oct 8, 2026
)

### What does this PR do?

Type of change: Refactor

### Summary of the series

Model-specific PTQ modeling moves out of
`modelopt/torch/quantization/plugins/huggingface.py` into the per-model
packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`,
next to each model's `specs.py`. By the end of the series
`huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the
code shared across models: sequential/fused MoE auto-detection,
attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the
transposed-quantization helpers that gpt_oss and llama4 share. Model
types that gain PTQ modeling but had no `ModelSpec` get one.

- Class and function bodies move verbatim. The only rewrites: on-the-fly
callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and
`is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily.
- `huggingface.py` imports every `modeling_ptq` from an explicit list.
The list sits after the generic wrappers these modules build on and
before the homogeneous decoder discoverer is registered, so Nemotron-H's
more specific discoverer still matches first.
- Each package's `__init__.py` keeps importing only `specs`, so `import
modelopt.torch.models` does not pull in quantization or transformers.
- No public API or quantization behavior changes; only private
(`_`-prefixed) names change module.

Merge in order. #2576 targets `main`; each later PR targets the one
before it.

1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and
step3p7
2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon
and Llama4 ← **this PR**
3. #2578 — Move GPT-OSS
4. #2580 — Move the Step family (step3p5, shared by step3p7)

The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there
was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was
dropped: #2670 removed its pre-5.12 wrapper.

### This PR [2/4]

Introduces `<model_type>/modeling_ptq.py` and moves the first three
models there:

- `nemotron_h`: `is_nemotron_h_model` / `get_nemotron_h_decoder_layers`
and their layerwise-calibration registration.
- `falcon`: the `FalconLinear` registration and
`register_falcon_linears_on_the_fly`.
- `llama4`: `_QuantLlama4TextExperts` and its registration.

The registrations move in the form #2670 gave them (a top-level
transformers import plus `if ... not in QuantModuleRegistry`), and
`huggingface.py` drops the imports it no longer uses. It also adds the
`importlib` loop in `huggingface.py` that imports each `modeling_ptq`,
documents the convention in `modelopt/torch/models/README.md` and the
package docstring, and adds
`tests/unit/torch/models/test_modeling_ptq_registration.py`, which
checks in a fresh interpreter per case that registration does not depend
on which module is imported first.

It also raises the `llama4`, `qwen3_vl_moe` and `falcon` specs from
#2576 to the transformers 5.5 floor that #2670 set for every other spec,
and drops the `qwen3_vl_moe` spec's pointer to the pre-5.12 wrapper
#2670 removed.

### Usage

No API change. To add PTQ support for a new model, put its wrapper in
`modelopt/torch/models/<model_type>/modeling_ptq.py` and add
`<model_type>` to the import loop at the end of
`quantization/plugins/huggingface.py`.

### Testing

Each branch of the stack checked out and tested on `main` at 90ba9fb,
CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models
tests/unit/torch/quantization tests/unit/torch/export
tests/unit/recipe`, 2370 passed, 1 skipped. With transformers 5.5 (the
new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and
`test_huggingface.py`, 205 passed, 1 skipped. GPU tests are left to CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Only private (`_`-prefixed)
names change module; public API and quantization behavior are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (verbatim move; covered by
existing tests)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal refactor)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Part of a 5-PR stack; see the series list above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added post-training quantization support for Llama 4 expert layers.
* Expanded Falcon quantization coverage to include linear layers
discovered at runtime and older remote-code checkpoints.
* Added Nemotron-H model detection and decoder-layer discovery for
quantization.

* **Bug Fixes**
* Improved Qwen3-VL mixture-of-experts compatibility by applying the
legacy quantization wrapper only when needed.

* **Documentation**
* Clarified how model-specific quantization support is loaded and where
model support information belongs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
@shengliangxu shengliangxu changed the title Move model-specific PTQ modeling into modelopt/torch/models [4/5] [OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [4/5] Oct 8, 2026
shengliangxu added a commit that referenced this pull request Oct 9, 2026
…dels [3/4] (#2578)

### What does this PR do?

Type of change: Refactor

### Summary of the series

Model-specific PTQ modeling moves out of
`modelopt/torch/quantization/plugins/huggingface.py` into the per-model
packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`,
next to each model's `specs.py`. By the end of the series
`huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the
code shared across models: sequential/fused MoE auto-detection,
attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the
transposed-quantization helpers that gpt_oss and llama4 share. Model
types that gain PTQ modeling but had no `ModelSpec` get one.

- Class and function bodies move verbatim. The only rewrites: on-the-fly
callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and
`is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily.
- `huggingface.py` imports every `modeling_ptq` from an explicit list.
The list sits after the generic wrappers these modules build on and
before the homogeneous decoder discoverer is registered, so Nemotron-H's
more specific discoverer still matches first.
- Each package's `__init__.py` keeps importing only `specs`, so `import
modelopt.torch.models` does not pull in quantization or transformers.
- No public API or quantization behavior changes; only private
(`_`-prefixed) names change module.

Merge in order. #2576 and #2577 have merged; #2578 now targets `main`,
and #2580 targets #2578's branch.

1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and
step3p7 (merged)
2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon
and Llama4 (merged)
3. #2578 — Move GPT-OSS  ← **this PR**
4. #2580 — Move the Step family (step3p5, shared by step3p7)

The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there
was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was
dropped: #2670 removed its pre-5.12 wrapper.

### This PR [3/4]

Moves `_QuantGptOssExperts` and its registration into
`gpt_oss/modeling_ptq.py`, adds `gpt_oss` to the import loop, and
extends the registration test to it. This slice originally also moved
the pre-5.12 Qwen3-VL-MoE wrapper; #2670 removed that wrapper, so that
half was dropped.

### Usage

No API change. To add PTQ support for a new model, put its wrapper in
`modelopt/torch/models/<model_type>/modeling_ptq.py` and add
`<model_type>` to the import loop at the end of
`quantization/plugins/huggingface.py`.

### Testing

Each branch of the stack checked out and tested on `main` at 90ba9fb,
CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models
tests/unit/torch/quantization tests/unit/torch/export
tests/unit/recipe`, 2371 passed, 1 skipped. With transformers 5.5 (the
new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and
`test_huggingface.py`, 206 passed, 1 skipped. GPU tests are left to CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Only private (`_`-prefixed)
names change module; public API and quantization behavior are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (verbatim move; existing
tests re-pointed)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal refactor)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Part of a 5-PR stack; see the series list above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added post-training quantization support for GPT-OSS models.
* Added support for quantizing legacy Qwen3-VL-MoE layouts, while
retaining support for newer fused layouts.
* Quantization support now accommodates both legacy and newer
Qwen3-VL-MoE model layouts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
shengliangxu added a commit that referenced this pull request Oct 9, 2026
…dels [4/4] (#2580)

### What does this PR do?

Type of change: Refactor

### Summary of the series

Model-specific PTQ modeling moves out of
`modelopt/torch/quantization/plugins/huggingface.py` into the per-model
packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`,
next to each model's `specs.py`. By the end of the series
`huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the
code shared across models: sequential/fused MoE auto-detection,
attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the
transposed-quantization helpers that gpt_oss and llama4 share. Model
types that gain PTQ modeling but had no `ModelSpec` get one.

- Class and function bodies move verbatim. The only rewrites: on-the-fly
callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and
`is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily.
- `huggingface.py` imports every `modeling_ptq` from an explicit list.
The list sits after the generic wrappers these modules build on and
before the homogeneous decoder discoverer is registered, so Nemotron-H's
more specific discoverer still matches first.
- Each package's `__init__.py` keeps importing only `specs`, so `import
modelopt.torch.models` does not pull in quantization or transformers.
- No public API or quantization behavior changes; only private
(`_`-prefixed) names change module.

Merge in order. #2576, #2577 and #2578 have merged; this last slice,
#2580, now targets `main`.

1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and
step3p7 (merged)
2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon
and Llama4 (merged)
3. #2578 — Move GPT-OSS (merged)
4. #2580 — Move the Step family (step3p5, shared by step3p7) ← **this
PR**

The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there
was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was
dropped: #2670 removed its pre-5.12 wrapper.

### This PR [4/4]

Moves `_QuantMoELinear`, `_is_expert_indexed_moe_linear`,
`_is_step_family_model`, `register_moe_linear_on_the_fly` and
`_reconstruct_fused_moe_linear` into `step3p5/modeling_ptq.py`, which
also serves `step3p7`. This carries over #2569's switch of
`_is_step_family_model` to `hf_model_type`; `huggingface.py` drops its
now-unused `hf_model_type` and `re` imports.

Export (`hf_export_handlers.py`, `layerwise_export.py`,
`unified_export_hf.py`, `unified_export_hf_streaming.py`) imports these
names from the new module. The imports stay deferred, because
`modelopt.torch.quantization` imports `modelopt.torch.export` and
importing at module scope would risk an import cycle; the comments now
give that reason. The PTQ skill references (`unsupported-models.md`,
`checkpoint-validation.md`) now point agents at
`modelopt/torch/models/<model_type>/modeling_ptq.py` for model-specific
patches.

Rebasing onto #2649 also replaces the README's DBRX example of a
model-specific wrapper with Llama4's fused BMM experts.

`step3p7` also gets its own `modeling_ptq.py`, which imports the
Step-3.5 module: Step-3.7 remote-code checkpoints use the same
`MoELinear`, and today they are covered only because the HF plugin
imports every `modeling_ptq`. With its own module, a future loader that
imports PTQ modeling by model type finds it too.

### Usage

No API change. To add PTQ support for a new model, put its wrapper in
`modelopt/torch/models/<model_type>/modeling_ptq.py` and add
`<model_type>` to the import loop at the end of
`quantization/plugins/huggingface.py`.

### Testing

Each branch of the stack checked out and tested on `main` at 90ba9fb,
CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models
tests/unit/torch/quantization tests/unit/torch/export
tests/unit/recipe`, 2372 passed, 1 skipped. With transformers 5.5 (the
new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and
`test_huggingface.py`, 207 passed, 1 skipped. GPU tests are left to CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Only private (`_`-prefixed)
names change module; public API and quantization behavior are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (verbatim move; existing
tests re-pointed)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal refactor)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Part of a 5-PR stack; see the series list above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added post-training quantization support for Step-family models with
expert-indexed MoE layers.
* Checkpoint exports preserve the original layout of quantized expert
weights and scales.
* **Documentation**
* Updated MoE architecture examples and guidance for configuring
model-specific quantization support, including Step model revisions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant