Skip to content

[OMNIML-5899] Add IQ unified checkpoint export - #2444

Closed
hychiang-git wants to merge 32 commits into
hungyuehc/omniml-5899-ggmlfrom
hungyuehc/omniml-5899-export
Closed

hychiang-git wants to merge 32 commits into
hungyuehc/omniml-5899-ggmlfrom
hungyuehc/omniml-5899-export

Conversation

@hychiang-git

@hychiang-git hychiang-git commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do?

Type of change: new feature

This is the export PR in a three-PR series. It:

  • identifies IQ1_S and IQ2_XS in the ModelOpt export path and emits quantization metadata for group size, packing, and payload bytes;
  • stores each quantized <module>.weight as a shaped uint8 payload in a unified Hugging Face checkpoint, without separate packed_weights or weight_shape keys;
  • handles standard linear weights and fused expert weights;
  • adds IQ1_S and IQ2_XS export for Hugging Face and tensor-parallel-size-1 Megatron models; and
  • documents the checkpoint layout and logical-shape recovery rule.

Megatron export rejects tensor parallelism greater than one with a collective early error until cross-rank packing is defined. The guard scans every enabled tensor quantizer, including mixed-format models. Packed Megatron expert layouts also raise an early error because their HF transpose moves the 256-value block axis; dense tensor-parallel-size-1 Megatron export remains supported.

Review follow-up:

  • enforces the tensor-parallel guard collectively before state gathering;
  • revalidates each owning module at the final packing boundary;
  • avoids false amax fallback warnings and mutations for fused IQ experts;
  • continues filtering legacy weight_shape metadata from older checkpoints;
  • centralizes IQ format metadata and packing dispatch;
  • rejects unsupported IQ search modes;
  • checks every packed payload against the format-defined shape, including leading dimensions;
  • reports every incompatible IQ weight shape before whole-model or layerwise packing mutates the model;
  • uses the shared IQ registry for format detection with explicit checkpoint-packer dispatch;
  • preserves IQ metadata in the Megatron fallback path; and
  • adds layerwise export coverage.

Further review follow-up:

  • rejects both unsupported packed-expert layouts before expert fan-in or transpose;
  • avoids retaining and stacking many expert weights on the accelerator; and
  • removes an unused field from the shared format specification;
  • applies the collective guard to extra-module export before gathering or packing;
  • scopes the guard to exporters that actually pack IQ weights, leaving fake-quant vLLM export unaffected; and
  • attributes both TypeError and ValueError packing failures while preserving their exception type;
  • validates every diffusion component before processing the first component;
  • rejects incompatible Megatron shapes collectively before main or extra-module state conversion; and
  • includes qualified weight labels in IQ configuration rejections;
  • propagates exact HF state-dict labels through dense and split-expert packing;
  • diagnoses unsupported formats before inspecting module quantizers; and
  • combines Megatron IQ presence and shape state in one collective while skipping the full shape scan for non-IQ models;
  • enumerates grouped source weights and permits plural fused-expert quantizer lists during preflight while rejecting unsupported singular custom weight attributes before mutation; and
  • attributes malformed packer payload shapes to the exact logical weight;
  • scans the union of grouped operation and quantizer counts so trailing expert weights cannot escape preflight; and
  • runs the model-wide IQ preflight from the shared checkpoint-preparation path used by direct callers;
  • removes the redundant top-level scan so the shared preparation guard remains the single source of truth; and
  • formats grouped expert prefixes before configuration checks so failures name a concrete expert; and
  • runs the IQ shape preflight in the separate accelerate-offload streaming setup before any shard is written.
  • separates the concrete IQ diagnostic label from the grouped-expert exclusion prefix; and
  • bypasses the detailed weight-shape scan when the model has no enabled IQ quantizer.
  • keeps the vLLM fake-quant Megatron override compatible with grouped expert diagnostic labels.
  • rejects packed-expert IQ layouts collectively before any pipeline rank starts conversion.
  • hoists IQ weight inspection into shared helpers and covers the legacy fused-expert quantizer layout.
  • identifies packed payload metadata as mandatory for consumers of the nominal bit-width fields.
  • collects unsupported IQ quantizer settings with shape failures so every pipeline rank exits in preflight.
  • describes current-device retention without implying that this branch already contains native packing.
  • runs shape and quantizer-setting preflight from every HF export entry point before weights or shards are mutated.
  • uses unprefixed activation-quantizer names for grouped experts while retaining concrete weight labels in diagnostics.

Usage

from modelopt.torch.export import export_hf_checkpoint

# `model` has already been quantized with an IQ recipe.
export_hf_checkpoint(model, export_dir="exported_model")

The exported weight shape is [*logical_shape[:-1], logical_shape[-1] // 256, payload_bytes], where payload_bytes is 50 for IQ1_S and 74 for IQ2_XS.

Testing

  • 158 tests passed in the combined local export and IQ suite before the latest Megatron-only guard tests.
  • The latest 184 relevant export, conversion, and layer-helper tests passed.
  • The latest export-weight and fused-expert checks covered 87 relevant cases: 86 passed together, and the corrected direct-entry regression passed separately.
  • A later combined run passed 85 unaffected cases; after two synthetic models were aligned with the real export contract, all 3 focused preflight regressions passed.
  • The offload-streaming preflight regression passed and verifies that failure leaves both the weight and export directory untouched.
  • The non-IQ fast-path regression passed; the grouped Megatron label and prefix regression compiles as Python.
  • The updated export module compiles as Python and all changed-file pre-commit hooks pass.
  • The latest reload/export suite passed 13 tests after rebasing onto the current core head.
  • The grouped vLLM fake-quant regression compiles as Python and passes all changed-file hooks; runtime remains for GPU CI because Megatron is unavailable locally.
  • The latest focused export run passed 31 IQ cases; the distributed Megatron regressions compile as Python.
  • The latest shared-preflight update passed 33 focused export cases.
  • Megatron and layerwise GPU coverage is included but was not run on the local macOS host; GPU CI remains required.

Before your PR is "Ready for review"

Make sure you read and follow Contributor guidelines and your commits are signed (git commit -s -S).

Make sure you read and follow the Security Best Practices (e.g. avoiding hardcoded trust_remote_code=True, torch.load(..., weights_only=False), pickle, etc.).

  • Is this change backward compatible?: ✅
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: N/A — no copied code or new PIP dependency.
  • Did you write any new necessary tests?: ✅
  • Did you update Changelog?: N/A — the series entry is in [OMNIML-5899] Add IQ1_S and IQ2_XS GGML quantization #2443.
  • Did you get Claude approval on this PR?: ❌ — a fresh review was requested after the current head update.

Additional Information

@copy-pr-bot

copy-pr-bot Bot commented Sep 16, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

@hychiang-git

Copy link
Copy Markdown
Contributor Author

/claude review

Comment thread modelopt/torch/export/unified_export_megatron.py Outdated
Comment thread modelopt/torch/export/moe_utils.py Outdated
Comment thread modelopt/torch/export/quant_utils.py
Comment thread modelopt/torch/export/convert_hf_config.py Outdated
Comment thread modelopt/torch/export/unified_export_hf.py Outdated
Comment thread modelopt/torch/export/quant_format.py
Comment thread modelopt/torch/export/quant_utils.py Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude review — IQ unified checkpoint export

Scope: full review (trigger comment carried no scoping instructions). 11 changed files (+343/-13); reviewed all of modelopt/torch/export/ (6 files) and all 4 test files, plus traced the new symbols into modelopt/torch/quantization/ggml/, layerwise_export.py, and quantization/utils/core_utils.py for callers. Only docs/source/deployment/3_unified_hf.rst was skimmed rather than reviewed in depth.

Findings: CRITICAL: 0 · IMPORTANT: 2 · SUGGESTION: 5

IMPORTANT

  1. TP>1 guard misses the second Megatron entry point (unified_export_megatron.py:316-324). The NotImplementedError sits in save_pretrained, but save_pretrained_extra_modules reads self.extra_state_dict, which reaches _get_iq_weight_state through the same _name_remapping chain. An EAGLE/Medusa + IQ export at TP>1 therefore packs each rank's partial weight into GGML blocks and writes it with no error. Moving the check into _get_iq_weight_state covers every entry point and cannot be bypassed by a future one.
  2. Fused-expert amax fallback now fires for every IQ projection (moe_utils.py:218 region). IQ quantizers are amax-free by construction, so the 'was not calibrated (amax missing or zero) … Consider using more calibration data' fallback is unconditionally true on this newly-enabled path. Because the message embeds the expert index, Python's duplicate-warning filter does not collapse it — a 128-expert model emits roughly 384 bogus warnings per MoE layer, and export writes a meaningless _amax onto each quantizer. Gate the block on the format needing an amax.

SUGGESTION (non-blocking, details inline)

  1. "weight_shape" added to weight_suffixes in postprocess_state_dict is dead — this design deliberately emits no shape companion, and a test asserts its absence.
  2. Block geometry (50/74/1.5625/2.3125/256) is hardcoded in three places while IQ1_S_BLOCK_BYTES, the EFFECTIVE_BITS constants, and GGML_BLOCK_SIZE already exist in quantization/ggml; also, the GGML-specific keys sit inside the compressed-tensors weights args rather than at the scheme level where this file's NVFP4_SVD precedent puts non-schema keys.
  3. Megatron's hand-built fallback quantization_config omits group_size/packing/block_payload_bytes for IQ, so the same format gets different hf_quant_config.json metadata depending on whether combined_layer_config_dict was populated.
  4. The qformat-to-packer dispatch is duplicated three times, and export ignores backend_extra_args['search_impl'] — harmless today (only 'auto' is accepted), but it means export is only guaranteed to match calibration by there being a single implementation.
  5. Adding IQ to FUSION_FREE_FORMATS also opts it into layerwise_export.SUPPORTED_FORMATS. That path traces as functional, but it is neither claimed in the PR body nor tested.

What checked out

  • The scale-free representation is coherent end to end: packed shape [*logical[:-1], logical[-1]//256, payload_bytes] makes the logical shape unambiguously recoverable, and the Megatron test's assert_close(dequantize(packed), weight_quantizer(weight)) verifies export bytes match what fake-quant produced — the property that actually matters here.
  • All seven weight-writing branches in unified_export_megatron.py (_name_remapping, gated MLP, expert shards, qkv slicing, GatedDeltaNet, and both merge paths) are patched, each with the IQ branch correctly ahead of weight_scale is None — without that ordering the weights would have silently exported as raw BF16. The GatedDeltaNet keep_bf16 interaction is handled.
  • get_quantization_format's string num_bits branch is safe: representative_weight_quantizer unwraps GroupedQuantizer to a TensorQuantizer, so .backend always resolves, and the branch sits after the SequentialQuantizer handling.
  • The moe_utils buffer-vs-parameter fix is necessary and correct — assigning a plain tensor to proj.weight would have become an ordinary attribute and dropped the packed payload from state_dict() entirely.
  • Mixed-precision config generation works for IQ via _quant_algo_to_group_config, and _postprocess_single_tensor passes the uint8 weight through untouched.

Risk: moderate. No correctness bug in the packing or metadata for the paths the PR claims. Finding 1 is a real silent-corruption hole, but only on the EAGLE/Medusa + IQ + TP>1 combination; finding 2 is a usability regression rather than a wrong checkpoint. Backward compatibility looks clean — every change is behind a new-format branch, no existing signature, default, or modelopt_state key moves. Since GPU CI has not run these paths, the Megatron coverage is still unproven in practice.

@hychiang-git
hychiang-git force-pushed the hungyuehc/omniml-5899-export branch from 756d86d to 28a6e1b Compare September 16, 2026 01:21
@hychiang-git

Copy link
Copy Markdown
Contributor Author

Addressed all seven review findings in 28a6e1bbf:

  • enforced the tensor-parallel guard in the shared state path;
  • skipped false amax fallback handling for fused IQ experts;
  • removed the obsolete shape special case;
  • centralized format metadata and packing dispatch;
  • rejected unsupported search modes;
  • preserved Megatron fallback metadata; and
  • added layerwise export coverage.

Local validation: 133 focused tests passed, and all pre-commit hooks passed. GPU-only coverage remains for CI.

@hychiang-git

Copy link
Copy Markdown
Contributor Author

/claude review

Comment thread modelopt/torch/export/unified_export_megatron.py Outdated
Comment thread modelopt/torch/export/unified_export_megatron.py Outdated
Comment thread modelopt/torch/export/unified_export_megatron.py
Comment thread modelopt/torch/export/quant_format.py Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude review — IQ unified checkpoint export (round 2, 28a6e1bb)

Scope: full review (trigger comment carried no scoping instructions). 12 changed files (+554/-16); reviewed all 6 modelopt/torch/export/ files and all 5 test files, plus re-traced into modelopt/torch/quantization/ggml/ (__init__.py, common.validate_weight, iq1_s, iq2_xs) and quantization/utils/core_utils.representative_weight_quantizer. docs/source/deployment/3_unified_hf.rst read but not deeply reviewed.

Findings: CRITICAL: 2 · IMPORTANT: 1 · SUGGESTION: 1

Prior round: all seven addressed, verified

  1. TP>1 guard now sits in _get_iq_weight_state (unified_export_megatron.py:1163), so it covers save_pretrained_extra_modules / EAGLE-Medusa too, plus a direct unit test — resolved.
  2. Fused-expert amax fallback is gated on is_iq at both the per-channel slice and the calibration fallback (moe_utils.py:172, :201), with a test asserting no warning and no _amax creation — resolved.
  3. "weight_shape" removed from _BASE_SKIP_KEYS; grep confirms nothing under modelopt/torch/ emits that key, so the removal is inert — resolved.
  4. Format metadata centralized in IQ_FORMAT_SPECS / iq_format_spec, sourced from the ggml constants; packing dispatch centralized in _pack_iq_weight — resolved.
  5. _validate_iq_quantizer_config rejects non-auto search_impl, mirroring the iq1_s.py:310 / iq2_xs.py:311 fake-quant check, with a test — resolved.
  6. Megatron fallback quantization_config now merges group_size/block_payload_bytes/packing (:406) — resolved.
  7. Layerwise coverage added (test_iq2_layerwise_export_matches_whole_model_export) — resolved.

I did not re-raise the config.json group-config metadata suggestion from last round: with quant_method: "modelopt" and quant_algo: "IQ1_S" both present in the emitted quantization_config, a consumer has a discriminator and won't mistake it for a plain compressed-tensors int1 scheme, so leaving packing/block_payload_bytes to hf_quant_config.json is a defensible call.

New this round

CRITICAL — Megatron packed-expert paths pack along the wrong axis (_pack_name_remapping:1920, _pack_name_remapping_gpt_oss:2033)

Both helpers stack per-expert weights to [E, out, in], transpose(-2, -1) into the HF [E, in, out] layout (and, for GPT-OSS, additionally interleave the gate/up halves along the last dim), and only then call _get_iq_weight_state. quantize_iq1_s/quantize_iq2_xs always reshape to (-1, 256), i.e. they block along the last dimension — which after the transpose is out, not the in reduction axis that iq*_fake_quant blocked along during calibration (block_sizes={-1: 256}).

So each 256-weight super-block covers a completely different element set than the one whose d and local scales were fitted. dequantize_iq*(exported) does not reproduce weight_quantizer(weight) — the same invariant the new dense test (test_megatron_name_remapping_exports_iq_payload) correctly asserts, and which no MoE test covers. It is also not GGML-interpretable (GGML blocks run along the GEMM reduction axis), and it can hard-fail validate_weight, whose divisibility-by-256 check applies to the post-transpose last dim (hidden_size for linear_fc2, 2 * ffn_hidden_size_per_partition for linear_fc1) rather than to in.

MoE Megatron export at TP=1 is inside the scope the PR body claims, and nothing guards it, so this currently produces a silently wrong checkpoint. Simplest correct move for this PR: raise NotImplementedError for IQ in both pack helpers, matching how TP>1 is handled, and define the packed-expert layout separately.

IMPORTANT — device-memory spike from keep_weight_device on the expert fan-in paths (:1113)

The flag is right for the single-module paths, but the two pack helpers call _get_quantized_state once per local expert and accumulate, so IQ now holds all num_local_experts bf16 weights on the accelerator and allocates two more E × out × in device copies (stack, then transpose().contiguous()) while the model's own expert weights are still live. Every other format does that fan-in on CPU. Multi-GB transient per layer on a 128-expert model, immediately before packing allocates the payload.

SUGGESTION — effective_bits in IQ_FORMAT_SPECS has no reader (details inline).

What checked out

  • The scale-free representation is coherent for the paths the dense tests cover: packed shape [*logical[:-1], logical[-1]//256, payload_bytes] makes the logical shape unambiguously recoverable given the divisibility precondition, and the Megatron test verifies dequantize(packed) == weight_quantizer(weight) rather than just shapes/dtypes — the property that actually matters.
  • _validate_iq_quantizer_config composes correctly with all three quantizer layouts: representative_weight_quantizer unwraps GroupedQuantizer (TEGroupedLinear) and the plural weight_quantizers ModuleList (_QuantFusedExperts), and the isinstance(..., TensorQuantizer) check correctly rejects a SequentialQuantizer rather than silently reading num_bits off the wrong object.
  • The moe_utils buffer-vs-parameter fix is necessary and correct: after _export_quantized_weight swaps wrapper.weight for a uint8 buffer, a plain proj.weight = tensor assignment would have become an ordinary attribute and dropped the payload from state_dict() entirely. The new test pins exactly that.
  • get_quantization_format's string-num_bits branch is safe — it sits after the SequentialQuantizer handling, and the backend != "ggml" rejection prevents an IQ name reaching a non-GGML packer.
  • HF-side _export_quantized_weight: weight_name != "weight" is rejected up front, so the hardcoded register_buffer("weight", ...) cannot diverge from weight_name; delattr before register_buffer is the right order for a Parameter → buffer swap.
  • Config generation: process_layer_quant_config produces the full IQ metadata for the uniform case, _quant_algo_to_group_config handles both the uniform and the MIXED_PRECISION per-group case (quant_algo.lower() correctly matching the upper-case "IQ1_S" stored in quantized_layers), and the Megatron fallback path now agrees with it.
  • Backward compatibility is clean: every change is behind a new-format branch or a defaulted keyword (keep_weight_device=False); no existing signature, default, or modelopt_state key moves. quant_format.py's new top-level import of modelopt.torch.quantization.ggml is import-safe — the CUDA extension is loaded lazily inside quantize_iq*, not at package import.
  • All seven unified_export_megatron weight-writing branches place the IQ check ahead of weight_scale is None; without that ordering the weights would have exported as raw BF16. The keep_bf16 interaction in the GatedDeltaNet path is handled.

Risk: moderate-to-high for MoE Megatron, low elsewhere. The dense HF and dense Megatron paths look correct and are now tested end-to-end on the property that matters. The packed-expert paths were wired in mechanically alongside them and are wrong in a way that produces a loadable-but-incorrect checkpoint with no error — that is the one thing to fix before merge. GPU and Megatron CI still have not run any of this.

@hychiang-git
hychiang-git force-pushed the hungyuehc/omniml-5899-export branch from 28a6e1b to 0098670 Compare September 16, 2026 04:29
@hychiang-git

Copy link
Copy Markdown
Contributor Author

Addressed all four round-two findings in 009867070:

  • both packed Megatron expert paths now raise before transpose or interleaving can change the calibrated block axis;
  • the early guard also prevents multi-expert accelerator retention and stacking;
  • dense tensor-parallel-size-1 Megatron export remains supported; and
  • the unused effective-bits entry was removed from the shared export specification.

Current local regression: 87 tests passed, and all pre-commit hooks passed. Megatron GPU coverage remains for CI.

@hychiang-git

Copy link
Copy Markdown
Contributor Author

/claude review

Comment thread tests/gpu_megatron/torch/export/test_unified_export_megatron.py Outdated
Comment thread modelopt/torch/export/quant_utils.py

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude review — IQ unified checkpoint export (round 3, 00986707)

Scope: full review (trigger comment carried no scoping instructions). 12 changed files (+584/-14); reviewed all 6 modelopt/torch/export/ files and all 5 test files, plus re-traced into modelopt/torch/quantization/ggml/ (__init__, common.validate_weight/validate_packed_weights, iq1_s.quantize_iq1_s, iq2_xs), export/plugins/mcore_custom.py, export/plugins/mcore_llama.py, and tests/gpu_megatron/conftest.py. docs/source/deployment/3_unified_hf.rst was read and its block-layout arithmetic checked, but not deeply reviewed.

Findings: CRITICAL: 0 · IMPORTANT: 1 · SUGGESTION: 1

Round-two findings: all four addressed, verified

  1. Packed-expert axis bug — both _pack_name_remapping (:1882) and _pack_name_remapping_gpt_oss (:1956) now raise NotImplementedError for IQ. The raise sits after _get_quantized_state for the first expert but before weight_list.append, torch.stack, transpose, the GPT-OSS interleave, _record_layer_quant_config, and any self._state_dict write — so it fires on iteration 0 with no partial mutation, which the new test pins via assert exporter._state_dict == {}. Resolved.
  2. Device-memory spike on expert fan-in — the same early raise removes the accumulate-then-stack path entirely for IQ. I separately confirmed the remaining keep_weight_device=True consumer, the TEGroupedLinear expert loop (:1455–:1500), holds at most one expert weight on the accelerator per iteration: _get_iq_weight_state packs and returns .detach().cpu() inside the shard loop, and the name_to_value["weight"] device copy is popped immediately. Resolved.
  3. Dense TP=1 Megatron still supported — confirmed; only the two pack helpers and TP>1 are gated.
  4. Unused effective_bits — removed from IQ_FORMAT_SPECS; each remaining key (quant_algo, num_bits, group_size, block_payload_bytes, packing) has a reader in process_layer_quant_config, _quant_algo_to_group_config, or the Megatron fallback. Resolved.

I also re-verified the seven round-one fixes still hold (TP guard in _get_iq_weight_state, is_iq amax gating in moe_utils, weight_shape removal from _BASE_SKIP_KEYS, centralized IQ_FORMAT_SPECS/_pack_iq_weight, search_impl rejection, Megatron fallback metadata, layerwise coverage).

New this round

IMPORTANT — test_megatron_name_remapping_exports_iq_payload errors instead of asserting (test_unified_export_megatron.py:113)

The test runs in the main pytest process and never patches get_tensor_model_parallel_world_size, which _get_iq_weight_state now calls unconditionally. tests/gpu_megatron/conftest.py only initializes model-parallel state inside spawned dist_workers, so there is no TP group and MCore asserts tensor model parallel group is not initialized. The two sibling TP tests you added patch this exact symbol for the same reason.

This is the only test that checks dequantize(packed) == weight_quantizer(weight) — the invariant that actually proves the Megatron dense path encodes along the calibrated axis. It will error in CI rather than validate the round-trip, and the shape/dtype/layer_config_dict assertions after it never run. One-line fix in the inline comment.

SUGGESTION — _validate_iq_quantizer_config does not reject an enabled input_quantizer/pre_quant_scale, and both IQ write paths return before activation metadata is emitted, so an IQ-weights + quantized-activations config would export a silently weight-only checkpoint (details inline).

What checked out

  • Block axis is preserved on every enabled path. I enumerated all seven weight-writing branches and confirmed each hands _get_iq_weight_state a tensor whose last dim is still the reduction axis: _name_remapping writes [out, in] unchanged; the gated-MLP split (:1289), the TEGroupedLinear gated shard split (:1483), and the GatedDeltaNet projections (:1778) all slice dim 0; _qkv_slicing reshapes to [qkv_dim, head_size, hidden_size] and _take restores [-1, hidden_size]. A grep for transpose/permute/view in the exporter turns up layout changes only in the two now-guarded pack helpers and _merge_nvfp4_expert_scales. So quantize_iq*'s reshape(-1, 256) blocks along the same axis block_sizes={-1: 256} calibrated on, and validate_weight's divisibility check applies to the real reduction dim.
  • The HF fused-MoE path is also axis-correct, which matters because it is enabled rather than guarded: first_proj is [E, 2*expert_dim, hidden], so first_proj[idx, :expert_dim, :] and down[idx] both keep in last — consistent with what the 3-D fake-quant blocked along.
  • _pack_name_remapping's transpose=False mode has no caller (mcore_llama.py:72/:76 are the only two PackNameRemapping sites, neither passes it), so the unconditional IQ raise there rejects nothing that would have been correct.
  • search_impl validation matches calibration exactly — extra_args.get("search_impl", extra_args.get("iq_search_impl", "auto")) is identical to the fake-quant checks at iq1_s.py:309 and iq2_xs.py:310, including the legacy alias, so export cannot accept an impl the quantizer rejected.
  • moe_utils buffer-vs-parameter fix is necessary and correct: after _export_quantized_weight swaps wrapper.weight for a uint8 buffer, proj.weight = tensor would have become a plain attribute and dropped the payload from state_dict(). The new test pins exactly that.
  • Config generation is self-consistent across all three producers — process_layer_quant_config (uniform + MIXED_PRECISION per-group), the Megatron fallback at :406, and _quant_algo_to_group_config. quant_algo.lower() in IQ_FORMATS correctly matches the upper-case "IQ1_S" stored in quantized_layers, iq_format_spec lowercases its argument, and the emitted quant_method: "modelopt" plus quant_algo: "IQ1_S" give a consumer a discriminator so the weights-only int group is not mistaken for plain compressed-tensors int1.
  • postprocess_state_dict passes the packed weight through untouched — "weight" matches no _BASE_SKIP_KEYS entry, so it takes the early continue, and _maybe_squeeze_scale is scale-keyed.
  • Docs arithmetic is right: 2 + 64 + 8 = 74 bytes, 32 groups / 16 shared local scales = 2 groups per scale, 2^9 = 512 codebook rows, 74*8/256 = 2.3125 bits, and the recovery rule [*shape[:-2], shape[-2]*256] inverts the packed shape given the divisibility precondition.
  • Import safety and backward compat: quant_format.py's new top-level modelopt.torch.quantization.ggml import is safe (the CUDA extension loads lazily inside quantize_iq*), and convert_hf_config.py's new dependency on it only reaches consumers already inside the torch export stack. Every behavioral change is behind a new-format branch or a defaulted keyword (keep_weight_device=False); no existing signature, default, or modelopt_state key moves.

Risk: low. The correctness picture is materially better than round two — the axis hazard is now gated on both packed-expert helpers with no partial-write window, and I could not find another enabled path where the exported blocks diverge from what fake-quant produced. The one remaining item is test-side but consequential: the single test proving the Megatron round-trip cannot run as written, so after fixing it please confirm Megatron CI actually goes green on this file rather than green-by-collection.

🤖 Generated with Claude Code

@hychiang-git
hychiang-git force-pushed the hungyuehc/omniml-5899-export branch from 0098670 to ea560c0 Compare September 16, 2026 04:52
@hychiang-git

Copy link
Copy Markdown
Contributor Author

Addressed both round-three findings in ea560c0eb:

  • patched the dense Megatron test's TP world-size lookup so it reaches the payload and reconstruction assertions; and
  • rejected enabled input quantization and active pre_quant_scale for IQ export, preventing a weight-only checkpoint from silently dropping activation metadata. Added coverage for both formats and both invalid modes.

Validation: 109/109 targeted export and IQ tests passed, and all changed-file pre-commit hooks passed. Megatron runtime coverage remains for GPU CI.

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
@hychiang-git
hychiang-git force-pushed the hungyuehc/omniml-5899-export branch from 3c474fe to d9be3bb Compare September 16, 2026 15:02
@hychiang-git

Copy link
Copy Markdown
Contributor Author

/claude review

Comment thread modelopt/torch/export/quant_utils.py
Comment thread modelopt/torch/export/quant_utils.py

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude review — IQ unified checkpoint export (head d9be3bb0, "Collect all IQ export preflight failures")

Findings: CRITICAL: 0 · IMPORTANT: 1 · SUGGESTION: 1

Scope

Bare /claude review, so a full pass. 22 changed files (+1724/-39). All 10 changed modelopt/torch/export/ files reviewed hunk-by-hunk with surrounding context; CHANGELOG.rst, docs/source/deployment/3_unified_hf.rst, and the modelopt_recipes/ diffs read in full; test diffs read from the diff (GPU / Megatron suites not executed). Traced into modelopt/torch/quantization/ggml/ (common.validate_weight / validate_packed_weights, iq1_s._quantize_iq1_s_packed / dequantize_iq1_s), quantization/utils/core_utils.py (representative_weight_quantizer, weight_attr_names, quantizer_attr_names, is_quantized_linear), nn/modules/tensor_quantizer.GroupedQuantizer, and export/plugins/mcore_custom.py + mcore_llama.py + mcore_gptoss.py for the packed-expert rule names.

Findings

Severity Location Finding
IMPORTANT quant_utils.py:195 The HF-side preflight collects only shape errors. _iq_export_quantizer_config_errors — the model-wide config collector this PR adds — is wired only into the Megatron guard, so search_impl != "auto", unknown backend_extra_args, and enabled activation quantizers / a live pre_quant_scale still raise mid-walk from inside _pack_iq_weight. Whole-model HF export leaves layers 0..N-1 already replaced by uint8 payloads in place; layerwise export leaves already-finished layer shards written under layerwise.export_dir. One-line fix using the collector that already exists.
SUGGESTION quant_utils.py:133 The GroupedQuantizer branch yields a weight<index> name, so quantizer_attr_names derives weight0_input_quantizer — an attribute a TEGroupedLinear does not have. The weight-only activation guard is a silent no-op for grouped experts in preflight, while pack time (weight_name="weight") does catch it.

Prior-round finding: addressed, verified

The round-6 suggestion — the packed-expert NotImplementedError in _pack_name_remapping / _pack_name_remapping_gpt_oss being the one non-collective IQ guard — is now closed. _custom_mapping_to_lambda stamps _modelopt_export_func_name onto the wrapped rule (lambda to named def, closure semantics unchanged), _model_has_packed_expert_iq_quantizer reduces self.rules to linear_fc1 / linear_fc2 for the PackNameRemapping / PackNameRemappingGPT entries in mcore_llama.py:72-79 and mcore_gptoss.py:41-48, and matches only modules with local_experts in their path — so a dense mlp.linear_fc1 cannot false-positive and a MoE arch whose rules use plain NameRemapping yields an empty set. self.rules = self.all_rules[self.arch] (:207, re-bound at :221 / :232 for Medusa) keeps the scan arch-scoped. The flag is folded into the existing MAX all-reduce, so every rank raises together before layer_state_dicts materialization.

What checked out this round

  • Packing contract. _pack_iq_weight's expected_shape still matches _quantize_iq1_s_packed's own packed_shape exactly, and IQ1_S_BLOCK_SIZE / IQ2_XS_BLOCK_SIZE both resolve to GGML_BLOCK_SIZE = 256, so the documented [*shape[:-2], shape[-2] * 256] recovery rule agrees with validate_packed_weights. The RuntimeError for a shape mismatch sits outside the except (TypeError, ValueError) attribution handler, so it is not relabeled; validate_weight's TypeError (non-float weight) and ValueError (empty / non-divisible / non-finite) both keep their type through raise type(exc)(...).
  • Preflight axis. weight.shape[-1] % group_size remains the packed axis for every path that packs: _gated_mlp_slicing (weight[:ffn_hidden_size]), _qkv_slicing (_take(...).reshape(-1, hidden_size)), _gated_delta_net_slicing (torch.split(..., dim=0)), _grouped_mlp_slicing (weight[:half]), HF fused experts (first_proj[idx, :expert_dim, :], down[idx]). The two transposing paths are the ones rejected.
  • No unguarded weight write. All seven _get_quantized_state call sites accounted for: five have an if qformat in IQ_FORMATS branch, two raise. Every state-dict weight write at :1342, :1398-1399, :1610, :1893 is inside an else reachable only for non-IQ, and _gated_delta_net_slicing's keep_bf16 projections .cpu() the split view explicitly now that the source stays on device.
  • keep_weight_device aliasing. _get_weight_bias(keep_weight_device=True) can return module.weight itself when dtype already matches, but every IQ consumer either packs into a fresh uint8 tensor (.detach().cpu()) or .cpu()-copies the keep_bf16 slice, and _name_remapping's weight = weight + 1.0 is out-of-place — so no state-dict entry aliases a live parameter.
  • EP gather. IQ payloads reach all_gather_object through the existing torch.save-bytes round trip (:1647-1666) that was added precisely because pickling uint8 quantized tensors trips on UntypedStorage, so the new format needs no extra handling there.
  • Config emission. _get_quantization_format(self.model) is get_quantization_format(model) with no dependence on layer_config_dict, so hoisting it above self.layer_state_dicts in save_pretrained is behavior-preserving. convert_hf_quant_config_format's new IQ_FORMATS branch reuses _quant_algo_to_group_config, and process_layer_quant_config emits the same four keys for the MIXED_PRECISION path.
  • Backward compatibility. module_name, describe_as, keep_weight_device are all optional with prior behavior as the default; VllmFqGPTModelExporter accepts describe_as for signature compatibility and opts out via _packs_iq_weights = False. No mode registration, modelopt_state schema, or public __init__.py change.
  • Docs. The 50-byte IQ1_S (2 + 32 + 16) and 74-byte IQ2_XS (2 + 64 + 8) field breakdowns match the llama.cpp block definitions, and 1.5625 / 2.3125 bits per weight are the correct bytes * 8 / 256. The modelopt_recipes/ptq.md rewrite now states the divisibility requirement for packed export rather than the old "export added separately".

Risk

Low. Every IQ path is gated on num_bits in IQ_FORMATS, and no shared branch changes behavior for existing formats. Unsupported configurations fail loudly with a qualified weight label. The one blocking item is not a wrong-checkpoint risk — it is that the "reject before we mutate anything" invariant the rest of this PR establishes is only enforced for shapes on the HF side, not for quantizer configs, and the layerwise variant of that gap can leave partial shards on disk. Residual coverage risk is unchanged: the Megatron and layerwise IQ suites still need GPU CI as the real gate.

🤖 Generated with Claude Code

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
@hychiang-git

Copy link
Copy Markdown
Contributor Author

/claude review

return f"{qualified_name}: nonstandard weight attributes are not supported"

weight = getattr(module, weight_name, None)
if not isinstance(weight, torch.Tensor):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[SUGGESTION] The shape preflight silently passes when the weight attribute it synthesized does not resolve to a tensor, which is the one case where the guard's "fail before any mutation" contract quietly stops holding.

For the GroupedQuantizer branch the weight name is derived, not observed: _iq_export_weights builds f"weight{index}" over range(max(num_gemms, len(grouped_quantizer))) (:137-138). If num_gemms overshoots the materialized parameters, or a future grouped-linear variant names its per-GEMM parameters anything other than weight0..weightN-1, getattr(module, weight_name, None) returns None, this early return None fires, and the divisibility check for that weight never runs. The config check still runs (it never touches the weight), so nothing looks wrong — the model passes preflight and then _get_iq_weight_state → _pack_iq_weight raises mid-walk, after earlier modules have already been packed in place. That is exactly the failure mode _validate_iq_export_support exists to prevent, so the silent skip is worth making visible.

The weight_attr_names branch is unaffected (those names come from named_parameters, so the attribute always exists) — only the synthesized grouped names can miss.

Suggested fix: keep the return None for the observed-name path, but have the grouped branch of _iq_export_weights skip indices whose parameter does not exist rather than yielding them, and log at debug level when a grouped module yields fewer weights than max(num_gemms, len(grouped_quantizer)). A cheap alternative that costs nothing at runtime is to distinguish "no such attribute" from "not a tensor" here and return an error string for the former, since a grouped module that reports num_gemms = N but has no weightK is a layout ModelOpt cannot pack anyway.

Comment on lines +313 to +332
def _model_has_packed_expert_iq_quantizer(self) -> bool:
"""Return whether this rank would route an IQ expert through a packed-expert rule."""
packed_projection_names = {
rule_name.rsplit(".", 1)[-1]
for rule_name, rule in getattr(self, "rules", {}).items()
if getattr(rule, "_modelopt_export_func_name", None)
in {"pack_name_remapping", "pack_name_remapping_gpt_oss"}
}
if not packed_projection_names:
return False

for module_name, module in self.model.named_modules():
path = module_name.split(".")
if (
"local_experts" in path
and path[-1] in packed_projection_names
and self._model_has_iq_quantizer(module)
):
return True
return False

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[SUGGESTION] The packed-expert guard re-derives the dispatch decision through two indirect signals — a private marker attribute stuffed onto the rule closure, plus a module-name heuristic — when the export walk already has an authoritative flag for it.

The actual dispatch at :817 is if not self.rules.get("use_packed_local_experts", False): — that single boolean, not the rule's func_name, is what decides whether local_experts.linear_fc1 is called with the whole local_experts list (packed) or with one expert.linear_fc1 at a time. The guard here instead reconstructs the answer from _modelopt_export_func_name on apply_mapping and then matches "local_experts" in path and path[-1] in packed_projection_names. Today the two agree — every arch with PackNameRemapping{,GPT} on local_experts.* also sets use_packed_local_experts: True (mcore_llama.py:71,188, mcore_gptoss.py:40,66) — but they are independent tables, so a future arch that sets the flag while pointing the rule at a differently-named func, or one that packs experts stored under a container not literally named local_experts, would leave the guard silently returning False. The consequence is not a wrong checkpoint (_pack_name_remapping still raises at :1996/:2070), but the raise would then be non-collective and mid-walk, which is precisely the hang this collective preflight was added to avoid.

Reading the flag directly removes the marker plumbing added to _custom_mapping_to_lambda (:1122-1128) and keeps the guard tied to the same table the dispatch reads:

def _model_has_packed_expert_iq_quantizer(self) -> bool:
    """Return whether this rank would route an IQ expert through a packed-expert rule."""
    if not getattr(self, "rules", {}).get("use_packed_local_experts", False):
        return False
    return any(
        "local_experts" in module_name.split(".") and self._model_has_iq_quantizer(module)
        for module_name, module in self.model.named_modules()
    )

If you prefer to keep the marker (it does document which projections pack, which the flag does not), consider asserting the two agree — i.e. that packed_projection_names is non-empty exactly when use_packed_local_experts is set — so a future divergence surfaces as a loud error rather than as a guard that stops guarding.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude review — IQ unified checkpoint export (ea612a87)

Findings: CRITICAL: 0 · IMPORTANT: 0 · SUGGESTION: 2 → approving; neither suggestion blocks.

Scope. Full pass (no scoping instructions in the trigger). 22 changed files (+1797/-39): all 11 changed modelopt/torch/export/ files read hunk-by-hunk with surrounding context, plus CHANGELOG, docs/source/deployment/3_unified_hf.rst, and the modelopt_recipes/ changes. Traced outward into quantization/utils/core_utils.py (weight_attr_names, representative_weight_quantizer, quantizer_attr_names), tensor_quantizer.GroupedQuantizer, ggml/common.cached_reconstruction, the plugins/mcore_custom.py / mcore_llama.py / mcore_gptoss.py mapping tables, and the shipped configs/numerics/iq1_s.yaml plus configs/ptq/presets/model/iq1_s.yaml. GPU-only suites were not executed.

Note: the local base tip has moved ahead of GitHub's merge base, so git diff also surfaces an iq2_xs.py:_encode_blocks line this PR does not own. Excluded — a three-way merge keeps the base version.

The two suggestions (details inline):

  1. _iq_export_weight_shape_error silently skips a weight it cannot resolve. Reachable only via the GroupedQuantizer branch, where the name is synthesized as weight<index> over max(num_gemms, len(grouped_quantizer)) instead of observed from named_parameters. A num_gemms overshoot or a differently-named grouped parameter passes preflight, then fails inside _pack_iq_weight mid-walk — the exact mutate-then-raise ordering this guard exists to prevent.
  2. _model_has_packed_expert_iq_quantizer re-derives a decision the walk already owns, via a private _modelopt_export_func_name marker on the rule closure plus a local_experts-in-path heuristic, while the dispatch at :817 keys off use_packed_local_experts in self.rules. They agree today; if they diverge the guard degrades to the non-collective raise at :1996/:2070 — back to the mid-walk NCCL-hang shape it was added to fix.

What I verified:

  • Packing contract / block axis. expected_shape is sourced from the IQ1_S and IQ2_XS block-size/byte constants via IQ_FORMAT_SPECS, so the layout cannot drift from the encoder. I independently enumerated all 7 _get_quantized_state callers in the Megatron file (:1323, :1370, :1560, :1680, :1842, :1993, :2067) — every one has an IQ branch ahead of elif weight_scale is None, so no site can write raw BF16 under an IQ quant_algo. Every split feeding a packer is along a leading dim, so weight.shape[-1] (what preflight checks) is the axis the packer blocks on. The two layout-changing helpers are the rejected ones.
  • HF fused experts. Confirmed the wrapper normalizes to [E, out, in] (fused_dim0 = first_proj.shape[1]), so first_proj[idx, :expert_dim, :] and down[idx] keep the reduction dim last — matching the 3-D fake quant and matching what preflight checks on the unsplit parameter. Both is_iq gates suppress the shared-weight_scale_2 and per-projection amax fallbacks, so no bogus _amax and no per-expert warning storm.
  • Preflight across all four quantizer layouts that weight_attr_names / representative_weight_quantizer support: standard weight; singular custom attr (BMM-style → correctly rejected as nonstandard, consistent with _export_quantized_weight's own raise); plural _QuantFusedExperts lists (allowed — rebuilt as standard per-expert wrappers); GroupedQuantizer (allowed; it is an nn.ModuleList, not a TensorQuantizer, hence its own branch). The legacy _first_proj_attr sentinel fallback resolves to a quantizer whose sibling input quantizer does exist.
  • Weight-only enforcement cannot false-positive. The preset's base_disable_all wildcard covers *output_quantizer too, and KV quantizers are *_bmm_quantizer on attention, so IQ plus FP8 KV cache still exports. The backend_extra_args unknown-key check matches the recipe exactly (only search_impl is set), and the non-auto rejection mirrors the calibration-side check.
  • Config emission agrees across all three producers (process_layer_quant_config, the Megatron fallback at :493, _quant_algo_to_group_config). The IQ check in get_quantization_format sits ahead of the 4-bit and 8-bit branches, so an 8-bit input_quantizer cannot mislabel an IQ linear as int8_sq; the earlier SequentialQuantizer branch cannot swallow it either. The weight-only config_groups entry matches the W4A16_AWQ / W8A16 precedent.
  • State-dict hygiene. _reconstruction_cache is a plain dict attribute, not a buffer, so a quantizer left attached after the IQ early return cannot leak a cached tensor into the checkpoint. _maybe_squeeze_scale is scale-key-gated, so a 3-D payload with leading dim 1 is not squeezed (which would break the documented shape recovery). No blanket dtype cast over the state dict; requires_grad=False means a later .to(bf16) skips the uint8 payload.
  • Distributed / mode-state. _collective_iq_export_flags MAX-reduces on a backend-appropriate device and handles the peer-owns-the-bad-shape case; the TP check inside _get_iq_weight_state backstops direct state_dict callers; _packs_iq_weights = False keeps the vLLM fake-quant exporter out. No mode registration, no modelopt_state schema change, no public API change. quant_format.py's new import is the in-tree ggml package (no optional extra, no cycle, CUDA extension still lazy), so no import_plugin() gate is needed.
  • FUSION_FREE_FORMATS: checked all three consumers — suppresses the multi-module fusion path, suppresses the per-layer forward (which is what keeps a uint8 weight out of a GEMM), and opts IQ into layerwise_export.SUPPORTED_FORMATS.
  • Backward compat: every new parameter is optional and defaults to prior behavior (module_name, describe_as, keep_weight_device=False). Hoisting _get_quantization_format above layer_state_dicts is order-safe (pure inspection, no collective). The apply_mapping rewrite is behaviorally identical to the previous lambda.
  • Prior-round items hold: the module-local deepcopy import fixes the test-patch scope; the CHANGELOG layerwise bullet now lists IQ1_S and IQ2_XS; the output_quantizer guard, the IQ_FORMATS unification across write branches, and removal of the dead nn.Parameter isinstance fork are all present. Docs arithmetic re-derived: 2+32+16 = 50 and 2+64+8 = 74 bytes; 11-bit IQ1_S index gives 2048 rows, 9-bit IQ2_XS gives 512; 1.5625 and 2.3125 bits per weight; the documented recovery rule inverts the packed shape exactly.

Risk: low. Additive throughout — every IQ path is gated on the format set, and no shared branch changes behavior for an existing format. Unsupported configurations (TP greater than 1, packed experts, nonstandard weight attrs, last dim not divisible by 256, a non-auto search impl, any activation quantization or live pre_quant_scale) fail loudly with a qualified weight label, and the shape preflight is wired into all four packing entry points so a rejection leaves the model unmutated.

Residual risk is coverage, not correctness: per the PR body the Megatron and layerwise IQ paths were not run locally, so GPU CI is the real gate — please confirm tests/gpu_megatron/torch/export/test_unified_export_megatron.py and tests/gpu/torch/export/test_layerwise_export.py go green on hardware rather than green-by-collection.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants