Skip to content

why is int8 slower than fp16? #60

Description

@BobHo5474

Thank you for providing such a simple quantization method, but now I have encountered a problem.
I used "modelopt.onnx.quantization" for int8 quantization and then tested the inference speed with trtexec.
In the end, I found that the inference speed of int8 was slower than fp16, regardless of whether I used minxmax or entropy for calibration.

The average inference speed of the model using fp16 before quantization on the RTX4090 was 6.74 ms, while after quantization, the model using int8 had an average inference speed of 8.62 ms.

I'm using TensorRT 8.6.1 and modelopt 0.15.1.

You can download the original model, calibration data, and execution scripts from the link.

Activity

  1. self-assigned this
    on Aug 24, 2024
  2. ajrasane commented on Mar 20, 2026

    @ajrasane
    Contributor

    Hello, I quantized the model shared above with the following command:

    python -m modelopt.onnx.quantization \
    	--onnx_path=/home/scratch.arasane_hw/models/origin.onnx \
    	--quantize_mode=int8 \
    	--calibration_method=max \
    	--high_precision_dtype=fp32 \
    	--op_types_to_quantize=Conv
    

    There is some performance gain, but it is lower than expected. The accurate QDQ placement will need to be discussed with the TensorRT team. Hence I would recommend raising an issue there.

    We recently introduced another feature that in modelopt called autotune. With autotune, the latency is significantly higher. This is the command I used:

    python -m modelopt.onnx.quantization \
    	--onnx_path=/home/scratch.arasane_hw/models/origin.onnx \
    	--quantize_mode=int8 \
    	--calibration_method=max \
    	--high_precision_dtype=fp32 \
    	--autotune
    

    I hope this resolves your issue.

    Modelopt version: latest main
    TensorRT version: 10.13.2.2

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions