Skip to content

Remove LinearActivationQuantizedTensor and all related code - #4258

Merged
jerryzh168 merged 4 commits into
mainfrom
gh/jerryzh168/82/head
Apr 10, 2026
Merged

jerryzh168 merged 4 commits into
mainfrom
gh/jerryzh168/82/head

Conversation

@jerryzh168

@jerryzh168 jerryzh168 commented Apr 9, 2026 •

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Summary:
Delete LinearActivationQuantizedTensor (and its alias to_linear_activation_quantized)
as part of the ongoing AQT cleanup. This tensor subclass wrapped a weight tensor to
intercept F.linear and quantize activations on the fly. It is no longer needed as
activation quantization is handled through other mechanisms.

Removed:

  • torchao/quantization/linear_activation_quantized_tensor.py (class definition)
  • Imports and __all__ exports in torchao/quantization/__init__.py
  • isinstance guard in _is_linear() in quant_api.py
  • isinstance branch in _quantization_type() in utils.py

Test Plan:

  • Verified no remaining imports or usage of LinearActivationQuantizedTensor or
    to_linear_activation_quantized in the codebase (only string literals in error
    messages of unrelated files remain)
  • CI should pass since no tests depended on this class

Summary:
Remove the calibration_flow tutorial folder (static_quant.py, gptq_like.py,
awq_like.py) which depended on LinearActivationQuantizedTensor and other
deprecated AQT APIs. Update documentation references that pointed to these
deleted files.

Test Plan:
- Verified no remaining references to calibration_flow/, static_quant.py,
  gptq_like.py, or awq_like.py in the codebase
- Doc links updated to point to remaining valid resources

[ghstack-poisoned]
Summary:
Delete `LinearActivationQuantizedTensor` (and its alias `to_linear_activation_quantized`)
as part of the ongoing AQT cleanup. This tensor subclass wrapped a weight tensor to
intercept `F.linear` and quantize activations on the fly. It is no longer needed as
activation quantization is handled through other mechanisms.

Removed:
- `torchao/quantization/linear_activation_quantized_tensor.py` (class definition)
- Imports and `__all__` exports in `torchao/quantization/__init__.py`
- `isinstance` guard in `_is_linear()` in `quant_api.py`
- `isinstance` branch in `_quantization_type()` in `utils.py`

Test Plan:
- Verified no remaining imports or usage of `LinearActivationQuantizedTensor` or
  `to_linear_activation_quantized` in the codebase (only string literals in error
  messages of unrelated files remain)
- CI should pass since no tests depended on this class

[ghstack-poisoned]
@pytorch-bot

pytorch-bot Bot commented Apr 9, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/ao/4258

Note: Links to docs will display an error until the docs builds have been completed.

❗ 1 Active SEVs

There are 1 currently active SEVs. If your PR is affected, please view them below:

⏳ No Failures, 7 Pending

As of commit c793f08 with merge base fee43b5 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

jerryzh168 added a commit that referenced this pull request Apr 9, 2026
Summary:
Delete `LinearActivationQuantizedTensor` (and its alias `to_linear_activation_quantized`)
as part of the ongoing AQT cleanup. This tensor subclass wrapped a weight tensor to
intercept `F.linear` and quantize activations on the fly. It is no longer needed as
activation quantization is handled through other mechanisms.

Removed:
- `torchao/quantization/linear_activation_quantized_tensor.py` (class definition)
- Imports and `__all__` exports in `torchao/quantization/__init__.py`
- `isinstance` guard in `_is_linear()` in `quant_api.py`
- `isinstance` branch in `_quantization_type()` in `utils.py`

Test Plan:
- Verified no remaining imports or usage of `LinearActivationQuantizedTensor` or
  `to_linear_activation_quantized` in the codebase (only string literals in error
  messages of unrelated files remain)
- CI should pass since no tests depended on this class

ghstack-source-id: 7979b7f
Pull Request resolved: #4258
@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Apr 9, 2026
@jerryzh168 jerryzh168 added the module: not user facing Use this tag if you don't want this PR to show up in release notes label Apr 9, 2026
@jerryzh168

Copy link
Copy Markdown
Contributor Author

@claude review

@claude

claude Bot commented Apr 9, 2026 •

Copy link
Copy Markdown

Claude finished @jerryzh168's task in 1m 23s —— View job


Review: Remove LinearActivationQuantizedTensor and all related code

  • Read the PR diff and changed files
  • Check for remaining references to removed code
  • Review correctness and completeness
  • Post review feedback

Verdict: This PR looks good. The deletion is clean and well-scoped.

Changes reviewed

  1. torchao/quantization/linear_activation_quantized_tensor.py (deleted) — Entire 307-line file removed. This was the class definition, all aten dispatch implementations, and the to_linear_activation_quantized alias.

  2. torchao/quantization/__init__.py — Import and __all__ entries for both LinearActivationQuantizedTensor and to_linear_activation_quantized removed cleanly.

  3. torchao/quantization/quant_api.py — Import removed and the isinstance(mod.weight, LinearActivationQuantizedTensor) guard removed from _is_linear() (line 183). The remaining guards (_AffineFakeQuantizedTensor, NonDynamicallyQuantizableLinear) still prevent double-quantization of shared weights per the TODO comment at line 180.

  4. torchao/quantization/utils.py — Import and the isinstance branch in _quantization_type() removed. The function falls through correctly to the hasattr(weight, "_quantization_type") check.

Minor nit: stale string references

There are a few leftover string literals referencing LinearActivationQuantizedTensor in other files. These are just comments/error messages and don't cause any functional issues, but cleaning them up would be nice for consistency:

  • torchao/quantization/linear_activation_scale.py:98 — error message says "LinearActivationQuantizedTensor: No specialized dispatch found for linear op" but this file defines WeightTensorWithLinearActivationScaleMetadata
  • torchao/quantization/linear_activation_scale.py:122 — comment says "Allow a model with LinearActivationQuantizedTensor weights..." but registers WeightTensorWithLinearActivationScaleMetadata
  • torchao/quantization/weight_tensor_linear_activation_quantization.py:165 — same stale error message
  • torchao/quantization/weight_tensor_linear_activation_quantization.py:204 — same stale comment
  • torchao/quantization/linear_activation_weight_observed_tensor.py:156 — same stale comment

These look like they were copy-pasted from the original LinearActivationQuantizedTensor when creating the replacement classes and were never updated. Not blocking, but a follow-up to fix these strings would improve code clarity. Fix stale string references →

No issues found

  • No remaining imports or runtime references to the deleted class
  • No test files reference LinearActivationQuantizedTensor
  • The _is_linear guard removal is safe — the class no longer exists so nothing can be an instance of it
  • The _quantization_type branch removal is safe for the same reason

LGTM.

… all related code"

Summary:
Delete `LinearActivationQuantizedTensor` (and its alias `to_linear_activation_quantized`)
as part of the ongoing AQT cleanup. This tensor subclass wrapped a weight tensor to
intercept `F.linear` and quantize activations on the fly. It is no longer needed as
activation quantization is handled through other mechanisms.

Removed:
- `torchao/quantization/linear_activation_quantized_tensor.py` (class definition)
- Imports and `__all__` exports in `torchao/quantization/__init__.py`
- `isinstance` guard in `_is_linear()` in `quant_api.py`
- `isinstance` branch in `_quantization_type()` in `utils.py`

Test Plan:
- Verified no remaining imports or usage of `LinearActivationQuantizedTensor` or
  `to_linear_activation_quantized` in the codebase (only string literals in error
  messages of unrelated files remain)
- CI should pass since no tests depended on this class

[ghstack-poisoned]
Summary:
Delete `LinearActivationQuantizedTensor` (and its alias `to_linear_activation_quantized`)
as part of the ongoing AQT cleanup. This tensor subclass wrapped a weight tensor to
intercept `F.linear` and quantize activations on the fly. It is no longer needed as
activation quantization is handled through other mechanisms.

Removed:
- `torchao/quantization/linear_activation_quantized_tensor.py` (class definition)
- Imports and `__all__` exports in `torchao/quantization/__init__.py`
- `isinstance` guard in `_is_linear()` in `quant_api.py`
- `isinstance` branch in `_quantization_type()` in `utils.py`

Test Plan:
- Verified no remaining imports or usage of `LinearActivationQuantizedTensor` or
  `to_linear_activation_quantized` in the codebase (only string literals in error
  messages of unrelated files remain)
- CI should pass since no tests depended on this class

[ghstack-poisoned]
jerryzh168 added a commit that referenced this pull request Apr 9, 2026
Summary:
Delete `LinearActivationQuantizedTensor` (and its alias `to_linear_activation_quantized`)
as part of the ongoing AQT cleanup. This tensor subclass wrapped a weight tensor to
intercept `F.linear` and quantize activations on the fly. It is no longer needed as
activation quantization is handled through other mechanisms.

Removed:
- `torchao/quantization/linear_activation_quantized_tensor.py` (class definition)
- Imports and `__all__` exports in `torchao/quantization/__init__.py`
- `isinstance` guard in `_is_linear()` in `quant_api.py`
- `isinstance` branch in `_quantization_type()` in `utils.py`

Test Plan:
- Verified no remaining imports or usage of `LinearActivationQuantizedTensor` or
  `to_linear_activation_quantized` in the codebase (only string literals in error
  messages of unrelated files remain)
- CI should pass since no tests depended on this class

ghstack-source-id: d60e636
Pull Request resolved: #4258
@jerryzh168

Copy link
Copy Markdown
Contributor Author

we'll address minor issues by removing these subclass in follow up PRs

@jerryzh168
jerryzh168 changed the base branch from gh/jerryzh168/82/base to main April 9, 2026 22:46
@jerryzh168
jerryzh168 merged commit 4c82ea7 into main Apr 10, 2026
35 of 36 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: not user facing Use this tag if you don't want this PR to show up in release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants