Repository navigation
[OMNIML-3817] Move model-specific PTQ modeling into modelopt/torch/models [4/5] - #2579
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (2)
🚧 Files skipped from review as they are similar to previous changes (1)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review. 📝 WalkthroughWalkthroughDBRX PTQ wrappers and dynamic expert registration now reside in a DBRX-specific module. The Hugging Face plugin imports that module through its model-type loop. The guide and specs comment point to the new implementation path. ChangesDBRX PTQ
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Refactor Merge Risk: ⚪ Minimal · up to The DBRX PTQ relocation has no identified behavior change that needs fixing before merge. Normal validation can proceed. 🚥 Pre-merge checks | ✅ 5 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (5 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## shengliangx/move-modeling_3 #2579 +/- ##
===============================================================
+ Coverage 76.16% 78.42% +2.25%
===============================================================
Files 635 636 +1
Lines 69702 69709 +7
===============================================================
+ Hits 53090 54670 +1580
+ Misses 16612 15039 -1573
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
a989c02 to
c21bc33
Compare
|
82fe8d0 to
cc9c1c2
Compare
c21bc33 to
28d7ec6
Compare
cc9c1c2 to
a1f35f4
Compare
28d7ec6 to
820d490
Compare
a1f35f4 to
561aa83
Compare
aae8551 to
9a18923
Compare
44fdb8d to
f29052e
Compare
9a18923 to
f29052e
Compare
|
This slice is dropped from the series: DBRX support was removed in #2649, so there is no DBRX PTQ modeling left to move. Its branch now equals |
) ### What does this PR do? Type of change: Refactor ### Summary of the series Model-specific PTQ modeling moves out of `modelopt/torch/quantization/plugins/huggingface.py` into the per-model packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`, next to each model's `specs.py`. By the end of the series `huggingface.py` shrinks from ~2025 to ~1525 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had no `ModelSpec` get one. - Class and function bodies move verbatim. The only rewrites: on-the-fly callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and `is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily. - `huggingface.py` imports every `modeling_ptq` from an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first. - Each package's `__init__.py` keeps importing only `specs`, so `import modelopt.torch.models` does not pull in quantization or transformers. - No public API or quantization behavior changes; only private (`_`-prefixed) names change module. Merge in order. #2576 targets `main`; each later PR targets the one before it. 1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and step3p7 ← **this PR** 2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon and Llama4 3. #2578 — Move GPT-OSS and Qwen3-VL-MoE 4. #2580 — Move the Step family (step3p5, shared by step3p7) The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there was nothing left to move. ### This PR [1/4] Registers a `ModelSpec` for each model type that gains `modeling_ptq.py` later in the series but has no spec yet. Each spec records only facts checked against transformers 4.57 and 5.14: - `llama4` and `qwen3_vl_moe`: fused MoE layout (`gate_up_proj`/`down_proj`), no gate/up pair, grouped export off. This is data only: both blocks were already detected as MoE (structurally and by name), and their expert containers are exported without name lookups. - `falcon`: dense, no sections. - `step3p5` / `step3p7`: `modeling_source="remote_code"` and intentionally no `MoESpec`. Step's expert projections sit directly on the MoE MLP with no `experts` container. Declaring the block would make `is_moe` claim it and send AWQ export into `get_experts_list`, which does not support that layout. `step3p7` gets its own package, following the `gemma4_text` / `gemma4` precedent. The exhaustive spec tables in `tests/unit/torch/models/test_model_specs.py` (`EXPECTED_MOE_LAYOUTS`, `root_class_names`) gain the two MoE rows. ### Usage No API change. The new specs are read through the existing registry, e.g. `get_spec("llama4").moe_spec`. ### Testing Each branch of the stack checked out and tested with `pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe` on CPU (torch 2.11, transformers 5.14), on `main` at 9f902ae (#2649): 2248 passed, 7 skipped. The new block names resolve in `test_specs_vs_transformers.py`, and the step3p5/step3p7 remote-code absence checks pass. Class names were also checked against the transformers 4.57 wheel. GPU tests and the transformers 4.57 matrix are left to CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Only private (`_`-prefixed) names change module; public API and quantization behavior are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ Spec tables extended (`EXPECTED_MOE_LAYOUTS`, `root_class_names`). - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (internal refactor) - Did you get Claude approval on this PR?: ❌ ### Additional Information Part of a 5-PR stack; see the series list above. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added model support for Falcon, Llama 4, Qwen3-VL-MoE, Step-3.5, and Step-3.7. * Llama 4 and Qwen3-VL-MoE support includes recognition of their fused expert layers. * Falcon, Llama 4, and Qwen3-VL-MoE require Transformers 4.57 or later. Step-3.5 and Step-3.7 use remote-code modeling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
) ### What does this PR do? Type of change: Refactor ### Summary of the series Model-specific PTQ modeling moves out of `modelopt/torch/quantization/plugins/huggingface.py` into the per-model packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`, next to each model's `specs.py`. By the end of the series `huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had no `ModelSpec` get one. - Class and function bodies move verbatim. The only rewrites: on-the-fly callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and `is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily. - `huggingface.py` imports every `modeling_ptq` from an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first. - Each package's `__init__.py` keeps importing only `specs`, so `import modelopt.torch.models` does not pull in quantization or transformers. - No public API or quantization behavior changes; only private (`_`-prefixed) names change module. Merge in order. #2576 targets `main`; each later PR targets the one before it. 1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and step3p7 2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon and Llama4 ← **this PR** 3. #2578 — Move GPT-OSS 4. #2580 — Move the Step family (step3p5, shared by step3p7) The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was dropped: #2670 removed its pre-5.12 wrapper. ### This PR [2/4] Introduces `<model_type>/modeling_ptq.py` and moves the first three models there: - `nemotron_h`: `is_nemotron_h_model` / `get_nemotron_h_decoder_layers` and their layerwise-calibration registration. - `falcon`: the `FalconLinear` registration and `register_falcon_linears_on_the_fly`. - `llama4`: `_QuantLlama4TextExperts` and its registration. The registrations move in the form #2670 gave them (a top-level transformers import plus `if ... not in QuantModuleRegistry`), and `huggingface.py` drops the imports it no longer uses. It also adds the `importlib` loop in `huggingface.py` that imports each `modeling_ptq`, documents the convention in `modelopt/torch/models/README.md` and the package docstring, and adds `tests/unit/torch/models/test_modeling_ptq_registration.py`, which checks in a fresh interpreter per case that registration does not depend on which module is imported first. It also raises the `llama4`, `qwen3_vl_moe` and `falcon` specs from #2576 to the transformers 5.5 floor that #2670 set for every other spec, and drops the `qwen3_vl_moe` spec's pointer to the pre-5.12 wrapper #2670 removed. ### Usage No API change. To add PTQ support for a new model, put its wrapper in `modelopt/torch/models/<model_type>/modeling_ptq.py` and add `<model_type>` to the import loop at the end of `quantization/plugins/huggingface.py`. ### Testing Each branch of the stack checked out and tested on `main` at 90ba9fb, CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe`, 2370 passed, 1 skipped. With transformers 5.5 (the new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and `test_huggingface.py`, 205 passed, 1 skipped. GPU tests are left to CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Only private (`_`-prefixed) names change module; public API and quantization behavior are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (verbatim move; covered by existing tests) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (internal refactor) - Did you get Claude approval on this PR?: ❌ ### Additional Information Part of a 5-PR stack; see the series list above. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added post-training quantization support for Llama 4 expert layers. * Expanded Falcon quantization coverage to include linear layers discovered at runtime and older remote-code checkpoints. * Added Nemotron-H model detection and decoder-layer discovery for quantization. * **Bug Fixes** * Improved Qwen3-VL mixture-of-experts compatibility by applying the legacy quantization wrapper only when needed. * **Documentation** * Clarified how model-specific quantization support is loaded and where model support information belongs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
…dels [3/4] (#2578) ### What does this PR do? Type of change: Refactor ### Summary of the series Model-specific PTQ modeling moves out of `modelopt/torch/quantization/plugins/huggingface.py` into the per-model packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`, next to each model's `specs.py`. By the end of the series `huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had no `ModelSpec` get one. - Class and function bodies move verbatim. The only rewrites: on-the-fly callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and `is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily. - `huggingface.py` imports every `modeling_ptq` from an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first. - Each package's `__init__.py` keeps importing only `specs`, so `import modelopt.torch.models` does not pull in quantization or transformers. - No public API or quantization behavior changes; only private (`_`-prefixed) names change module. Merge in order. #2576 and #2577 have merged; #2578 now targets `main`, and #2580 targets #2578's branch. 1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and step3p7 (merged) 2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon and Llama4 (merged) 3. #2578 — Move GPT-OSS ← **this PR** 4. #2580 — Move the Step family (step3p5, shared by step3p7) The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was dropped: #2670 removed its pre-5.12 wrapper. ### This PR [3/4] Moves `_QuantGptOssExperts` and its registration into `gpt_oss/modeling_ptq.py`, adds `gpt_oss` to the import loop, and extends the registration test to it. This slice originally also moved the pre-5.12 Qwen3-VL-MoE wrapper; #2670 removed that wrapper, so that half was dropped. ### Usage No API change. To add PTQ support for a new model, put its wrapper in `modelopt/torch/models/<model_type>/modeling_ptq.py` and add `<model_type>` to the import loop at the end of `quantization/plugins/huggingface.py`. ### Testing Each branch of the stack checked out and tested on `main` at 90ba9fb, CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe`, 2371 passed, 1 skipped. With transformers 5.5 (the new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and `test_huggingface.py`, 206 passed, 1 skipped. GPU tests are left to CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Only private (`_`-prefixed) names change module; public API and quantization behavior are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (verbatim move; existing tests re-pointed) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (internal refactor) - Did you get Claude approval on this PR?: ❌ ### Additional Information Part of a 5-PR stack; see the series list above. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added post-training quantization support for GPT-OSS models. * Added support for quantizing legacy Qwen3-VL-MoE layouts, while retaining support for newer fused layouts. * Quantization support now accommodates both legacy and newer Qwen3-VL-MoE model layouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
…dels [4/4] (#2580) ### What does this PR do? Type of change: Refactor ### Summary of the series Model-specific PTQ modeling moves out of `modelopt/torch/quantization/plugins/huggingface.py` into the per-model packages under `modelopt/torch/models/<model_type>/modeling_ptq.py`, next to each model's `specs.py`. By the end of the series `huggingface.py` shrinks from ~1880 to ~1500 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention, `FP8Linear`, `CompressedLinear`, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had no `ModelSpec` get one. - Class and function bodies move verbatim. The only rewrites: on-the-fly callbacks are added to `CUSTOM_MODEL_PLUGINS` by their own module, and `is_homogeneous_hf_model` imports `is_nemotron_h_model` lazily. - `huggingface.py` imports every `modeling_ptq` from an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first. - Each package's `__init__.py` keeps importing only `specs`, so `import modelopt.torch.models` does not pull in quantization or transformers. - No public API or quantization behavior changes; only private (`_`-prefixed) names change module. Merge in order. #2576, #2577 and #2578 have merged; this last slice, #2580, now targets `main`. 1. #2576 — Add model specs for llama4, qwen3_vl_moe, falcon, step3p5 and step3p7 (merged) 2. #2577 — Add the `modeling_ptq.py` convention; move Nemotron-H, Falcon and Llama4 (merged) 3. #2578 — Move GPT-OSS (merged) 4. #2580 — Move the Step family (step3p5, shared by step3p7) ← **this PR** The DBRX slice (#2579) was dropped: #2649 removed DBRX support, so there was nothing left to move. Likewise, the Qwen3-VL-MoE half of [3/4] was dropped: #2670 removed its pre-5.12 wrapper. ### This PR [4/4] Moves `_QuantMoELinear`, `_is_expert_indexed_moe_linear`, `_is_step_family_model`, `register_moe_linear_on_the_fly` and `_reconstruct_fused_moe_linear` into `step3p5/modeling_ptq.py`, which also serves `step3p7`. This carries over #2569's switch of `_is_step_family_model` to `hf_model_type`; `huggingface.py` drops its now-unused `hf_model_type` and `re` imports. Export (`hf_export_handlers.py`, `layerwise_export.py`, `unified_export_hf.py`, `unified_export_hf_streaming.py`) imports these names from the new module. The imports stay deferred, because `modelopt.torch.quantization` imports `modelopt.torch.export` and importing at module scope would risk an import cycle; the comments now give that reason. The PTQ skill references (`unsupported-models.md`, `checkpoint-validation.md`) now point agents at `modelopt/torch/models/<model_type>/modeling_ptq.py` for model-specific patches. Rebasing onto #2649 also replaces the README's DBRX example of a model-specific wrapper with Llama4's fused BMM experts. `step3p7` also gets its own `modeling_ptq.py`, which imports the Step-3.5 module: Step-3.7 remote-code checkpoints use the same `MoELinear`, and today they are covered only because the HF plugin imports every `modeling_ptq`. With its own module, a future loader that imports PTQ modeling by model type finds it too. ### Usage No API change. To add PTQ support for a new model, put its wrapper in `modelopt/torch/models/<model_type>/modeling_ptq.py` and add `<model_type>` to the import loop at the end of `quantization/plugins/huggingface.py`. ### Testing Each branch of the stack checked out and tested on `main` at 90ba9fb, CPU, torch 2.11. With transformers 5.18: `pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipe`, 2372 passed, 1 skipped. With transformers 5.5 (the new floor): `tests/unit/torch/models` plus `test_moe_linear.py` and `test_huggingface.py`, 207 passed, 1 skipped. GPU tests are left to CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Only private (`_`-prefixed) names change module; public API and quantization behavior are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (verbatim move; existing tests re-pointed) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (internal refactor) - Did you get Claude approval on this PR?: ❌ ### Additional Information Part of a 5-PR stack; see the series list above. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added post-training quantization support for Step-family models with expert-indexed MoE layers. * Checkpoint exports preserve the original layout of quantized expert weights and scales. * **Documentation** * Updated MoE architecture examples and guidance for configuring model-specific quantization support, including Step model revisions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
What does this PR do?
Type of change: Refactor
Summary of the series
Model-specific PTQ modeling moves out of
modelopt/torch/quantization/plugins/huggingface.pyinto the per-model packages undermodelopt/torch/models/<model_type>/modeling_ptq.py, next to each model'sspecs.py. By the end of the serieshuggingface.pyshrinks from ~2145 to ~1490 lines and keeps only the code shared across models: sequential/fused MoE auto-detection, attention,FP8Linear,CompressedLinear, parallel linears, and the transposed-quantization helpers that gpt_oss and llama4 share. Model types that gain PTQ modeling but had noModelSpecget one.CUSTOM_MODEL_PLUGINSby their own module, andis_homogeneous_hf_modelimportsis_nemotron_h_modellazily.huggingface.pyimports everymodeling_ptqfrom an explicit list. The list sits after the generic wrappers these modules build on and before the homogeneous decoder discoverer is registered, so Nemotron-H's more specific discoverer still matches first.__init__.pykeeps importing onlyspecs, soimport modelopt.torch.modelsdoes not pull in quantization or transformers._-prefixed) names change module.Merge in order. #2576 targets
main; each later PR targets the one before it.modeling_ptq.pyconvention; move Nemotron-H, Falcon and Llama4This PR [4/5]
Moves
_QuantDbrxExperts,_QuantDbrxExpertGLU,_QuantDbrxFFN, their registrations andregister_dbrx_moe_on_the_flyintodbrx/modeling_ptq.py. The DBRX comment indbrx/specs.pyand the customized-model quantization guide point at the new file.Usage
No API change. To add PTQ support for a new model, put its wrapper in
modelopt/torch/models/<model_type>/modeling_ptq.pyand add<model_type>to the import loop at the end ofquantization/plugins/huggingface.py.Testing
Each branch of the stack checked out and tested with
pytest tests/unit/torch/models tests/unit/torch/quantization tests/unit/torch/export tests/unit/recipeon CPU (torch 2.11, transformers 5.14), onmainat e3c903c (the #2569 squash): 2155 passed, 7 skipped, plus 1 failure intest_autoquant_shapley.py::test_data_parallel_aumann_shapley, a multi-process test that is flaky onmainitself (it failed 2 of 4 runs onmainwith none of these changes). The stack has since been rebased ontomainat 44bbd47; the commits in between do not touch code under this change (the last one only editsAGENTS.md). GPU tests and the transformers 4.57 matrix are left to CI.Before your PR is "Ready for review"
_-prefixed) names change module; public API and quantization behavior are unchanged.CONTRIBUTING.md: N/AAdditional Information
Part of a 5-PR stack; see the series list above.
🤖 Generated with Claude Code
Summary by CodeRabbit