Skip to content

Add RunningAbsMaxSmoothQuantObserver for memory-efficient calibration (#3946) - #3946

Merged
meta-codesync[bot] merged 2 commits into
pytorch:mainfrom
jcaip:export-D94260071
Apr 2, 2026
Merged

meta-codesync[bot] merged 2 commits into
pytorch:mainfrom
jcaip:export-D94260071

Conversation

@jcaip

@jcaip jcaip commented Feb 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary:

Add a memory-efficient SmoothQuant observer that uses running per-channel
absmax instead of storing all calibration inputs. This reduces calibration
memory from O(N x features) to O(features), preventing RAM spikes and OOM
kills when calibrating on large datasets.

  • Add RunningAbsMaxSmoothQuantObserver class in core.py
  • Add use_running_absmax config option to SmoothQuantConfig
  • Export the new observer from the module

Differential Revision: D94260071

@pytorch-bot

pytorch-bot Bot commented Feb 25, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/ao/3946

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 04581a1 with merge base 960f307 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Feb 25, 2026
@meta-codesync

meta-codesync Bot commented Feb 25, 2026

Copy link
Copy Markdown

@jcaip has exported this pull request. If you are a Meta employee, you can view the originating Diff in D94260071.

@@ -0,0 +1,249 @@
# Copyright (c) Meta Platforms, Inc. and affiliates.

@namgyu-youn namgyu-youn Mar 1, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure why test script is located inside package. Shouldn't we update https://github.com/pytorch/ao/blob/main/test/prototype/test_smoothquant.py instead?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, will tell Hossein to do that before merging

return smoothing_factor, None


class RunningAbsMaxSmoothQuantObserver(torch.nn.Module):

@namgyu-youn namgyu-youn Mar 1, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Like this new logic overally, but do we actually need RunningAbsMaxSmoothQuantObserver? To me, minimal SmoothQuantObserver update (even though its bc-breaking) looks better.

Also, RunningAbsMaxSmoothQuantObserver reminds me AQT (AffineQuantizedTensor) nightmare. I am not sure if we really need abstracted notation.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah agree SmoothQuantObserver abstraction is kind of funky.

I think the best way this mess is to define a ObserverTensor class that we can then modify appropriately. That way we don't need to have SmoothQuantObserver, SmoothQuantObservedLinear and can just use a single SmoothQuantObserverTensor.

Do you know if smoothquant API is used much anywhere? The only thing that concerns me is that by default replacing SmoothQuantObserver with RunningAbsMaxSmoothQuantObserver is that we need two passes through the data for RunningAbsMaxSmoothQuantObserver, so existing workflows might fail.

We do need a way to test numerics though between SmoothQuantObserver and RunningAbsMaxSmoothQuantObserver, so even if we make RunningAbsMaxSmoothQuantObserver we'll still need to store this code somewhere. I think it's better to leave the cleanup as a follow up.

@namgyu-youn namgyu-youn Mar 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the best way this mess is to define a ObserverTensor class that we can then modify appropriately. That way we don't need to have SmoothQuantObserver, SmoothQuantObservedLinear and can just use a single SmoothQuantObserverTensor.

So are you considering common observer class? More generally, quantization subclass (Float8Tensor/Int8Tensor) should not contain observer-related ops (like act_pre_scale)? Actually I was working for this topic, so it would be great if you can look into it — #3925.

@namgyu-youn namgyu-youn Mar 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you know if smoothquant API is used much anywhere? The only thing that concerns me is that by default replacing SmoothQuantObserver with RunningAbsMaxSmoothQuantObserver is that we need two passes through the data for RunningAbsMaxSmoothQuantObserver, so existing workflows might fail.

Yeah I am also interested in this topic. For example, because AWQ uses greedy search

for i in range(self.scale_options):
ratio = i * 1 / self.scale_options
scales = x_max.pow(ratio).to(self.weight.dtype).clamp(min=1e-4).view(-1)
if best_scales is None:
best_scales = torch.ones_like(scales)
scales = scales / (scales.max() * scales.min()).sqrt()
config_handler = _QUANTIZE_CONFIG_HANDLER[type(self.base_config)]
dummy_mod = DummyModule(self.weight * scales)
quant_mod = config_handler(dummy_mod, self.base_config)
w = quant_mod.weight
orig_out = F.linear(acc, self.weight, self.bias)
q_out = F.linear(acc / scales, w, self.bias)
loss = (orig_out - q_out).pow(2).mean().item()
if loss < best_loss:
best_scales = scales
best_loss = loss

, we can try better searching algorithm like binary/tenary. It's GPTQ, but I already observed tenary is 2x faster than greedy: ModelCloud/GPTQModel#2419

Wondering if you have any candidates for these numerical techniques, what about your thought?

@namgyu-youn namgyu-youn Mar 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We do need a way to test numerics though between SmoothQuantObserver and RunningAbsMaxSmoothQuantObserver, so even if we make RunningAbsMaxSmoothQuantObserver we'll still need to store this code somewhere. I think it's better to leave the cleanup as a follow up.

I think in this early stage, we don't need multiple-configs to store all of them. Also I don't want to see RunningAbsMaxSmoothQuantObserver anywhere.

Instead, we can just run perf using https://github.com/pytorch/ao/blob/main/benchmarks/quantization/calibration_based/measure_accuracy_and_performance.sh and update minimally across AWQ/GPTQ/SmoothQuant I feel.

…pytorch#3946)

Summary:

Add a memory-efficient SmoothQuant observer that uses running per-channel
absmax instead of storing all calibration inputs. This reduces calibration
memory from O(N x features) to O(features), preventing RAM spikes and OOM
kills when calibrating on large datasets.

- Add RunningAbsMaxSmoothQuantObserver class in core.py
- Add use_running_absmax config option to SmoothQuantConfig
- Export the new observer from the module

Differential Revision: D94260071
@meta-codesync meta-codesync Bot changed the title Add RunningAbsMaxSmoothQuantObserver for memory-efficient calibration Add RunningAbsMaxSmoothQuantObserver for memory-efficient calibration (#3946) Mar 16, 2026
@jcaip
jcaip force-pushed the export-D94260071 branch from 9104532 to f86e0aa Compare March 16, 2026 17:05
@jcaip jcaip added the module: inference quantize_ api inference flow label Mar 16, 2026
@jcaip

jcaip commented Mar 16, 2026

Copy link
Copy Markdown
Contributor Author

@claude can you fix lint please

@claude

claude Bot commented Mar 16, 2026 •

Copy link
Copy Markdown

Claude finished @jcaip's task in 10m 33s —— View job


Fixed lint issues

  • Identify linting tools and configuration — repo uses ruff (v0.11.6) for linting and formatting
  • Run linter to find issues — CI ruff format --check failed on torchao/quantization/quantize_/common/quantization_step.py
  • Fix formatting in quantization_step.py
  • Commit and push fixes

Changes made

Applied ruff format style fixes to torchao/quantization/quantize_/common/quantization_step.py:

  1. Collapsed single-line docstring closing """ onto the same line (e.g., """Insert observers before running calibration flow.""")
  2. Added blank line after each docstring expression (standard ruff format)
  3. Wrapped long enum value assignments that exceeded 88-char line limit with parentheses:
    • PREPARE_FOR_SMOOTHQUANT_SMOOTHING_FACTOR
    • PREPARE_FOR_SMOOTHQUANT_ACTIVATION_SCALES

Note: I was unable to install/run ruff locally due to permission restrictions, so I applied the formatting rules manually based on Black/ruff conventions. If CI still fails, the formatting may need further adjustment.
| Branch

@meta-codesync

meta-codesync Bot commented Mar 17, 2026

Copy link
Copy Markdown

@jcaip has imported this pull request. If you are a Meta employee, you can view this in D94260071.

@meta-codesync
meta-codesync Bot merged commit 85c853e into pytorch:main Apr 2, 2026
21 of 23 checks passed
@namgyu-youn

namgyu-youn commented Apr 2, 2026 •

Copy link
Copy Markdown
Contributor

@jcaip @jerryzh168 do you have any GPU profiling e2e results for this PR? If I remember correctly, there wasn't a profiling benchmark at the time — just wanted to check if that's been added or if there are any benchmark numbers available I am missing.

Priyjain-amd pushed a commit to Priyjain-amd/ao that referenced this pull request May 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. fb-exported meta-exported module: inference quantize_ api inference flow

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants