Repository navigation
make eval script also handle performance measurement - #3473
Conversation
|
Stack from ghstack (oldest at bottom): |
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/ao/3473
Note: Links to docs will display an error until the docs builds have been completed. ✅ You can merge normally! (1 Unrelated Failure)As of commit 5c8d110 with merge base 486fe0d ( BROKEN TRUNK - The following job failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: e1d713e ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: 15f7481 ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: 15f7481 ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: 665f2c8 ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: cae97ab ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: 42466df ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: 79c5722 ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: a0019d4 ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: 404c330 ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: e44305f ghstack-comment-id: 3634216524 Pull-Request: #3473
Summary: 1. refactors the eval script to also handle performance measurement in vllm 2. adds a simple `vllm bench latency` script to bench in vllm The script is broken on every single recipe, we'll have to fix and enable things in future PRs, will update the performance tables afterwards. Also, add convenience flags to skip model creation, lm_eval, vllm as needed to enable running just a single model + single step. Test Plan: ``` SKIP_MODEL_CREATE=1 SKIP_LM_EVAL=1 SKIP_VLLM=0 with-proxy ./benchmarks/quantization/measure_accuracy_and_performance.sh h100 ``` Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: a200c88 ghstack-comment-id: 3634216524 Pull-Request: #3473
Updating quantization README with updated benchmark results from #3473, and deleting the old outdated benchmarks
Summary: Since #3473 landed which benchmarks LLaMa models with HF model definition and vllm, we no longer need the previous benchmarks we had which used a custom model definition. Deleting all of them. Note that `benchmarks/_models/` is using HF model definition, but deleting them too as #3473 is a more modern version of these which includes performance. Test Plan: CI Reviewers: Subscribers: Tasks: Tags: ghstack-source-id: 1c71a53 ghstack-comment-id: 3693211304 Pull-Request: #3552
|
@vkuzo Thanks for working on this new benchmark module; this looks much better than the old versions. I have some questions about this work:
ao/.github/scripts/torchao_model_releases/quantize_and_upload.py Lines 739 to 808 in 415e0e8 |
I want the solution to the "string to config" problem to be fully decoupled from having good benchmarks, to not block progress on either. I went with something simple for this PR, if torchao converges to a "string to config" util that everyone is good with we can switch these benchmarks to use it. I would note that |
We should support them eventually! |
|
Hi @vkuzo , May I ask what's the plan of the future running scripts? For example, for LLama model, I can see the plan is:
From the May I ask the following questions?
My concern is that we could reach peak performance by doing some tricks like modifying how the model runs. Like
Thanks! |
The motivation for the recent changes is (a) modernize benchmarking for recent changes in the field (such as popularity of vllm) and (b) get away from custom model definitions in torchao (since they are not widely used). For LLMs, we want to benchmark in vLLM to ensure we use the latest and greatest in the LLM OSS community. We are also adding a diffusers benchmark in #3502.
I'd actually recommend the diffusers benchmark from #3502 once it lands, it's simpler than vllm because this is not an autoregressive model - no prefill vs decode, no k-v cache, etc.
Yes, and vLLM has all of those tricks, so if you run the vLLM benchmarks you will capture the tricks :)
Improvements to the scripts are welcome! Please feel free to submit a PR, or file an issue if you'd like help from torchao team. |
Is this something that worked in the old scripts but does not in the new ones? If so, let's just fix it in the new scripts - let me know more context. |
Hi @vkuzo , thanks for the reply!
I will also do some more detailed analysis. Thanks for the help and happy new year~! |
Updating quantization README with updated benchmark results from #3473, and deleting the old outdated benchmarks
|
Hi, @vkuzo, I have some questions. I would like to clarify whether measure_accuracy_and_performance.sh is considered a stable and recommended benchmarking method for TorchAO. Specifically, is this script intended to serve as a long-term reference for measuring both accuracy and performance across different TorchAO features and backends? In addition, we would like to understand whether other low-precision data types and quantization methods (e.g., AWQ, GPTQ) are evaluated using the same benchmarking methodology. If not, what are the key differences in evaluation setup, and how should results across different quantization approaches be compared in a consistent and fair manner? |
yes
#3602 adds a similar path for calibration based approaches |
|
@vkuzo, thanks for your reply! |
Summary:
vllm
vllm bench latencyscript to bench in vllm for prefill and decodeAlso, add convenience flags to skip model creation, lm_eval, vllm as
needed to enable running just a single model + single step.
Test Plan:
Results (on H100):
Results (on B200):
Reviewers:
Subscribers:
Tasks:
Tags: