Skip to content

make eval script also handle performance measurement - #3473

Merged
vkuzo merged 15 commits into
mainfrom
gh/vkuzo/183/head
Dec 26, 2025
Merged

vkuzo merged 15 commits into
mainfrom
gh/vkuzo/183/head

Conversation

@vkuzo

@vkuzo vkuzo commented Dec 9, 2025 •

Copy link
Copy Markdown
Contributor

Summary:

  1. refactors the eval script to also handle performance measurement in
    vllm
  2. adds a simple vllm bench latency script to bench in vllm for prefill and decode

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
// full output: https://www.internalfb.com/phabricator/paste/view/P2094641791

Results (on H100):

Library Versions:
================================================================================
torch.__version__: 2.9.0+cu128
torch.cuda.get_device_name(): NVIDIA H100
torchao.__version__: 0.14.0+git5c8a14207
vllm.__version__: 0.13.0

Quantization Recipe Results:
================================================================================
+--------------------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+
| Recipe                   |   Checkpoint |     Wikitext |   Winogrande |   Winogrande |   Prefill |   Decode |   Speedup |   Speedup |
|                          |         (GB) |   Perplexity |          Acc |       Stderr |    toks/s |   toks/s |   Prefill |    Decode |
+==========================+==============+==============+==============+==============+===========+==========+===========+===========+
| None                     |        16.08 |       7.5435 |       0.7419 |       0.0123 |   30946.5 |  6612    |     1     |     1     |
+--------------------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+
| float8_rowwise           |         9.1  |       7.5919 |       0.7348 |       0.0124 |   45312.5 |  8025.95 |     1.464 |     1.214 |
+--------------------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+
| int8_rowwise_weight_only |         9.11 |       7.5561 |       0.7427 |       0.0123 |   28231.9 |  4309.8  |     0.912 |     0.652 |
+--------------------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+
| int8_rowwise             |         9.1  |       7.6567 |       0.738  |       0.0124 |           |          |           |           |
+--------------------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+

Results (on B200):

torch.__version__: 2.9.0+cu128
torch.cuda.get_device_name(): NVIDIA B200
torchao.__version__: 0.16.0+git5ad41c539
vllm.__version__: 0.13.0

Quantization Recipe Results:
================================================================================
+----------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+
| Recipe         |   Checkpoint |     Wikitext |   Winogrande |   Winogrande |   Prefill |   Decode |   Speedup |   Speedup |
|                |         (GB) |   Perplexity |          Acc |       Stderr |    toks/s |   toks/s |   Prefill |    Decode |
+================+==============+==============+==============+==============+===========+==========+===========+===========+
| None           |        16.08 |       7.5435 |       0.7427 |       0.0123 |   59099.9 |  14380   |     1     |     1     |
+----------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+
| mxfp8          |         9.32 |       7.6034 |       0.7316 |       0.0125 |           |          |           |           |
+----------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+
| nvfp4          |         6.05 |       8.4459 |       0.7135 |       0.0127 |  102786   |  15218.9 |     1.739 |     1.058 |
+----------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+
| float8_rowwise |         9.1  |       7.6263 |       0.7348 |       0.0124 |   69313.7 |  15984   |     1.173 |     1.112 |
+----------------+--------------+--------------+--------------+--------------+-----------+----------+-----------+-----------+

Reviewers:

Subscribers:

Tasks:

Tags:

vkuzo added 4 commits December 9, 2025 06:30
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
@vkuzo

vkuzo commented Dec 9, 2025 •

Copy link
Copy Markdown
Contributor Author

Stack from ghstack (oldest at bottom):

@pytorch-bot

pytorch-bot Bot commented Dec 9, 2025 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/ao/3473

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 5c8d110 with merge base 486fe0d (image):

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

vkuzo added a commit that referenced this pull request Dec 9, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: e1d713e
ghstack-comment-id: 3634216524
Pull-Request: #3473
@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Dec 9, 2025
@vkuzo vkuzo added the topic: for developers Use this tag if this PR is mainly developer facing label Dec 9, 2025
[ghstack-poisoned]
[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 10, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: 15f7481
ghstack-comment-id: 3634216524
Pull-Request: #3473
[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 10, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: 15f7481
ghstack-comment-id: 3634216524
Pull-Request: #3473
@vkuzo
vkuzo changed the base branch from gh/vkuzo/182/head to main December 10, 2025 18:09
[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 23, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: 665f2c8
ghstack-comment-id: 3634216524
Pull-Request: #3473
[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 23, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: cae97ab
ghstack-comment-id: 3634216524
Pull-Request: #3473
[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 23, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: 42466df
ghstack-comment-id: 3634216524
Pull-Request: #3473
[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 23, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: 79c5722
ghstack-comment-id: 3634216524
Pull-Request: #3473
[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 23, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: a0019d4
ghstack-comment-id: 3634216524
Pull-Request: #3473
[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 23, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: 404c330
ghstack-comment-id: 3634216524
Pull-Request: #3473

@jainapurva jainapurva left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks!

[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 26, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: e44305f
ghstack-comment-id: 3634216524
Pull-Request: #3473
[ghstack-poisoned]
vkuzo added a commit that referenced this pull request Dec 26, 2025
Summary:

1. refactors the eval script to also handle performance measurement in
   vllm
2. adds a simple `vllm bench latency` script to bench in vllm

The script is broken on every single recipe, we'll have to fix and
enable things in future PRs, will update the performance tables
afterwards.

Also, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.

Test Plan:

```
SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100
```

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: a200c88
ghstack-comment-id: 3634216524
Pull-Request: #3473
@vkuzo
vkuzo merged commit 886ac8e into main Dec 26, 2025
54 of 59 checks passed
vkuzo added a commit that referenced this pull request Dec 26, 2025
Updating quantization README with updated benchmark results from #3473, and deleting the old outdated benchmarks
vkuzo added a commit that referenced this pull request Dec 26, 2025
Summary:

Since #3473 landed which benchmarks
LLaMa models with HF model definition and vllm, we no longer need the
previous benchmarks we had which used a custom model definition.
Deleting all of them.

Note that `benchmarks/_models/` is using HF model definition, but
deleting them too as #3473 is a more
modern version of these which includes performance.

Test Plan: CI

Reviewers:

Subscribers:

Tasks:

Tags:
ghstack-source-id: 1c71a53
ghstack-comment-id: 3693211304
Pull-Request: #3552
@namgyu-youn

namgyu-youn commented Dec 28, 2025 •

Copy link
Copy Markdown
Contributor

@vkuzo Thanks for working on this new benchmark module; this looks much better than the old versions. I have some questions about this work:

  1. How about using an academic notation like W8A16-INT or W8A8-INT, instead of int8_rowwise_weight_only or int8_rowwise? Abstractions like granularity or HQQ can be added at the end of them I feel.

  2. Is there a plan supporting calibration-based apis like AWQ, SmoothQuant, and GPTQ? Those configs might be:

awq_config = AWQConfig(base_config, step="prepare")
if safe_serialization:
quantize_(model, awq_config, filter_fn=filter_fn_skip_lmhead)
else:
quantize_(model, awq_config)
TransformerEvalWrapper(
model=model,
tokenizer=tokenizer,
max_seq_length=max_seq_length,
).run_eval(
tasks=tasks,
limit=calibration_limit,
)
awq_config = AWQConfig(base_config, step="convert")
if safe_serialization:
quantize_(model, awq_config, filter_fn=filter_fn_skip_lmhead)
else:
quantize_(model, awq_config)
quantized_model = model
quant_config = AWQConfig(base_config, step="prepare_for_loading")
if safe_serialization:
quantization_config = TorchAoConfig(quant_config).to_dict()
quantized_model.config.quantization_config = quantization_config
hf_quantizer, _, _, _ = get_hf_quantizer(
config=quantized_model.config,
quantization_config=None,
dtype=torch.bfloat16,
device_map="cuda:0",
weights_only=True,
user_agent={
"file_type": "model",
"framework": "pytorch",
"from_auto_class": False,
},
)
quantized_model.hf_quantizer = hf_quantizer
else:
quantized_model.config.quantization_config = TorchAoConfig(quant_config)
elif quant == "SmoothQuant-INT8-INT8":
model = AutoModelForCausalLM.from_pretrained(
model_to_quantize,
device_map="auto",
torch_dtype=torch.bfloat16,
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
base_config = Int8DynamicActivationInt8WeightConfig()
quant_config = SmoothQuantConfig(base_config, step="prepare")
quantize_(
model,
quant_config,
)
TransformerEvalWrapper(
model=model,
tokenizer=tokenizer,
max_seq_length=max_seq_length,
).run_eval(
tasks=tasks,
limit=calibration_limit,
)
quant_config = SmoothQuantConfig(base_config, step="convert")
quantize_(model, quant_config)
quantized_model = model
load_config = SmoothQuantConfig(base_config, step="prepare_for_loading")
quantized_model.config.quantization_config = TorchAoConfig(load_config)

@vkuzo

vkuzo commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor Author

How about using an academic notation like W8A16-INT or W8A8-INT, instead of int8_rowwise_weight_only or int8_rowwise? Abstractions like granularity or HQQ can be added at the end of them I feel.

I want the solution to the "string to config" problem to be fully decoupled from having good benchmarks, to not block progress on either. I went with something simple for this PR, if torchao converges to a "string to config" util that everyone is good with we can switch these benchmarks to use it.

I would note that W8A16-INT is ambiguous as it does not specify how exactly things are quantized, so we'll need something more specific.

@vkuzo

vkuzo commented Dec 29, 2025

Copy link
Copy Markdown
Contributor Author

Is there a plan supporting calibration-based apis like AWQ, SmoothQuant, and GPTQ?

We should support them eventually!

@Stonepia

Stonepia commented Dec 30, 2025 •

Copy link
Copy Markdown
Contributor

Hi @vkuzo , May I ask what's the plan of the future running scripts?

For example, for LLama model, I can see the plan is:

  1. benchmarks/_models/eval_hf_models.py <-- deleted
  2. benchmarks/quantization/measure_accuracy_and_performance.sh <-- Prefer this
  3. torchao/_models/llama/benchmarks.sh <--- Going to be deleted

From the measure_accuracy_and_performance.sh, we could only benchmark through vllm bench, and the accuracy check goes to lm_eval. Seems that the plan is, torchao not responsible for holding any running scripts, just a wrapper, torchao prefer to call tools to reduce efforts.

May I ask the following questions?

  1. What is the recommended one for doing the simple performance collection task?

My concern is that we could reach peak performance by doing some tricks like modifying how the model runs. Like torch.compile for example. But if we totally rely on the benchmarking with vllm bench, do we need to move all of these tricks inside vLLM?

  1. Current scripts do not take care of the memory constraint GPUs. It gets OOM easily on my local XPU machines on int4-wo models. This may be because a) lm_eval have a copy b) some potential bugs missing. In any way, if we don't keep one script on torchao side, we need to do debugging on lm_eval, is this correct?

Thanks!

@vkuzo

vkuzo commented Dec 30, 2025

Copy link
Copy Markdown
Contributor Author

Seems that the plan is, torchao not responsible for holding any running scripts, just a wrapper, torchao prefer to call tools to reduce efforts.

The motivation for the recent changes is (a) modernize benchmarking for recent changes in the field (such as popularity of vllm) and (b) get away from custom model definitions in torchao (since they are not widely used). For LLMs, we want to benchmark in vLLM to ensure we use the latest and greatest in the LLM OSS community. We are also adding a diffusers benchmark in #3502.

What is the recommended one for doing the simple performance collection task?

I'd actually recommend the diffusers benchmark from #3502 once it lands, it's simpler than vllm because this is not an autoregressive model - no prefill vs decode, no k-v cache, etc.

My concern is that we could reach peak performance by doing some tricks like modifying how the model runs.

Yes, and vLLM has all of those tricks, so if you run the vLLM benchmarks you will capture the tricks :)

Current scripts do not take care of the memory constraint GPUs. It gets OOM easily on my local XPU machines on int4-wo models.

Improvements to the scripts are welcome! Please feel free to submit a PR, or file an issue if you'd like help from torchao team.

@vkuzo

vkuzo commented Dec 30, 2025

Copy link
Copy Markdown
Contributor Author

It gets OOM easily on my local XPU machines on int4-wo models.

Is this something that worked in the old scripts but does not in the new ones? If so, let's just fix it in the new scripts - let me know more context.

@Stonepia

Copy link
Copy Markdown
Contributor

Is this something that worked in the old scripts but does not in the new ones? If so, let's just fix it in the new scripts - let me know more context.

Hi @vkuzo , thanks for the reply!
For the OOM, this is not the regression. Should be because of two reasons:

  1. The llama/generate.py first load the model, then to(GPU), then quant. This to(GPU) will cause OOM.
  2. The lm_eval seems maintain the duplicate copy of model (or maybe not delete the model on time), so on the small memory machine, it would cause OOM.

I will also do some more detailed analysis.

Thanks for the help and happy new year~!

vkuzo added a commit that referenced this pull request Jan 5, 2026
Updating quantization README with updated benchmark results from #3473, and deleting the old outdated benchmarks
@xiaowangintel

Copy link
Copy Markdown
Collaborator

Hi, @vkuzo, I have some questions.

I would like to clarify whether measure_accuracy_and_performance.sh is considered a stable and recommended benchmarking method for TorchAO.

Specifically, is this script intended to serve as a long-term reference for measuring both accuracy and performance across different TorchAO features and backends?

In addition, we would like to understand whether other low-precision data types and quantization methods (e.g., AWQ, GPTQ) are evaluated using the same benchmarking methodology. If not, what are the key differences in evaluation setup, and how should results across different quantization approaches be compared in a consistent and fair manner?

@vkuzo

vkuzo commented Jan 29, 2026

Copy link
Copy Markdown
Contributor Author

I would like to clarify whether measure_accuracy_and_performance.sh is considered a stable and recommended benchmarking method for TorchAO. Specifically, is this script intended to serve as a long-term reference for measuring both accuracy and performance across different TorchAO features and backends?

yes

In addition, we would like to understand whether other low-precision data types and quantization methods (e.g., AWQ, GPTQ) are evaluated using the same benchmarking methodology. If not, what are the key differences in evaluation setup, and how should results across different quantization approaches be compared in a consistent and fair manner?

#3602 adds a similar path for calibration based approaches

@xiaowangintel

Copy link
Copy Markdown
Collaborator

@vkuzo, thanks for your reply!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. topic: for developers Use this tag if this PR is mainly developer facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants