Thank you for providing such a simple quantization method, but now I have encountered a problem.
I used "modelopt.onnx.quantization" for int8 quantization and then tested the inference speed with trtexec.
In the end, I found that the inference speed of int8 was slower than fp16, regardless of whether I used minxmax or entropy for calibration.
The average inference speed of the model using fp16 before quantization on the RTX4090 was 6.74 ms, while after quantization, the model using int8 had an average inference speed of 8.62 ms.
I'm using TensorRT 8.6.1 and modelopt 0.15.1.
You can download the original model, calibration data, and execution scripts from the link.
Thank you for providing such a simple quantization method, but now I have encountered a problem.
I used "modelopt.onnx.quantization" for int8 quantization and then tested the inference speed with trtexec.
In the end, I found that the inference speed of int8 was slower than fp16, regardless of whether I used minxmax or entropy for calibration.
The average inference speed of the model using fp16 before quantization on the RTX4090 was 6.74 ms, while after quantization, the model using int8 had an average inference speed of 8.62 ms.
I'm using TensorRT 8.6.1 and modelopt 0.15.1.
You can download the original model, calibration data, and execution scripts from the link.