diff --git a/docs/tutorials/index.md b/docs/tutorials/index.md
index 32bdf8aba4cc..d0b427fdb78a 100644
--- a/docs/tutorials/index.md
+++ b/docs/tutorials/index.md
@@ -104,6 +104,8 @@ Select API:
* [Module to Gluon API](/tutorials/python/module_to_gluon.html)
* [Gluon end to end from training to inference](/tutorials/gluon/gluon_from_experiment_to_deployment.html)
* [Automatic Mixed Precision in Gluon](/tutorials/amp/amp_tutorial.html)
+ * [How to build and install MXNet with MKL-DNN backend](/tutorials/mkldnn/MKLDNN_README.html)
+ * [How to quantize custom models with MKL-DNN backend](/tutorials/mkldnn/mkldnn_quantization.html) (new!)
* API Guides
* Core APIs
* NDArray
@@ -156,7 +158,6 @@ Select API:
* [Large-Scale Multi-Host Multi-GPU Image Classification](/tutorials/vision/large_scale_classification.html)
* [Importing an ONNX model into MXNet](/tutorials/onnx/super_resolution.html)
* [Optimizing Deep Learning Computation Graphs with TensorRT](/tutorials/tensorrt/inference_with_trt.html)
- * [How to build and install MXNet with MKL-DNN backend](/tutorials/mkldnn/MKLDNN_README.html)
* API Guides
* Core APIs
* NDArray
diff --git a/docs/tutorials/mkldnn/mkldnn_quantization.md b/docs/tutorials/mkldnn/mkldnn_quantization.md
new file mode 100644
index 000000000000..459bf2a17d40
--- /dev/null
+++ b/docs/tutorials/mkldnn/mkldnn_quantization.md
@@ -0,0 +1,259 @@
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+# Quantize custom models with MKL-DNN backend
+
+This document is to introduce how to quantize the customer models from FP32 to INT8 with Apache/MXNet toolkit and APIs under Intel CPU.
+
+If you are not familiar with Apache/MXNet quantization flow, please reference [quantization blog](https://medium.com/apache-mxnet/model-quantization-for-production-level-neural-network-inference-f54462ebba05) first, and the performance data is shown in [Apache/MXNet C++ interface](https://github.com/apache/incubator-mxnet/tree/master/cpp-package/example/inference) and [GluonCV](https://gluon-cv.mxnet.io/build/examples_deployment/int8_inference.html).
+
+## Installation and Prerequisites
+
+Installing MXNet with MKLDNN backend is an easy and essential process. You can follow [How to build and install MXNet with MKL-DNN backend](https://mxnet.incubator.apache.org/tutorials/mkldnn/MKLDNN_README.html) to build and install MXNet from source. Also, you can install the release or nightly version via PyPi and pip directly by running:
+
+```
+# release version
+pip install mxnet-mkl
+# nightly version
+pip install mxnet-mkl --pre
+```
+
+## Image Classification Demo
+
+A quantization script [imagenet_gen_qsym_mkldnn.py](https://github.com/apache/incubator-mxnet/blob/master/example/quantization/imagenet_gen_qsym_mkldnn.py) has been designed to launch quantization for image-classification models. This script is integrated with [Gluon-CV modelzoo](https://gluon-cv.mxnet.io/model_zoo/classification.html), so that all pre-trained models can be downloaded from Gluon-CV and then converted for quantization. For details, you can refer [Model Quantization with Calibration Examples](https://github.com/apache/incubator-mxnet/blob/master/example/quantization/README.md).
+
+## Integrate Quantization Flow to Your Project
+
+Quantization flow works for both symbolic and Gluon models. If you're using Gluon, you can first refer [Saving and Loading Gluon Models](https://mxnet.incubator.apache.org/versions/master/tutorials/gluon/save_load_params.html) to hybridize your computation graph and export it as a symbol before running quantization.
+
+In general, the quantization flow includes 4 steps. The user can get the acceptable accuracy from step 1 to 3 with minimum effort. Most of thing in this stage is out-of-box and the data scientists and researchers only need to focus on how to represent data and layers in their model. After a quantized model is generated, you may want to deploy it online and the performance will be the next key point. Thus, step 4, calibration, can improve the performance a lot by reducing lots of runtime calculation.
+
+
+
+Now, we are going to take Gluon ResNet18 as an example to show how each step work.
+
+### Initialize Model
+
+```python
+import logging
+import mxnet as mx
+from mxnet.gluon.model_zoo import vision
+from mxnet.contrib.quantization import *
+
+logging.basicConfig()
+logger = logging.getLogger('logger')
+logger.setLevel(logging.INFO)
+
+batch_shape = (1, 3, 224, 224)
+resnet18 = vision.resnet18_v1(pretrained=True)
+resnet18.hybridize()
+resnet18.forward(mx.nd.zeros(batch_shape))
+resnet18.export('resnet18_v1')
+sym, arg_params, aux_params = mx.model.load_checkpoint('resnet18_v1', 0)
+# (optional) visualize float32 model
+mx.viz.plot_network(sym)
+```
+First, we download resnet18-v1 model from gluon modelzoo and export it as a symbol. You can visualize float32 model. Below is a raw residual block.
+
+
+
+#### Model Fusion
+
+```python
+sym = sym.get_backend_symbol('MKLDNN_QUANTIZE')
+# (optional) visualize fused float32 model
+mx.viz.plot_network(sym)
+```
+It's important to add this line to enable graph fusion before quantization to get better performance. Below is a fused residual block. Batchnorm, Activation and elemwise_add are fused into Convolution.
+
+
+
+### Quantize Model
+
+A python interface `quantize_graph` is provided for the user. Thus, it is very flexible for the data scientist to construct the expected models based on different requirements in a real deployment.
+
+```python
+# quantize configs
+# set exclude layers
+excluded_names = []
+# set calib mode.
+calib_mode = 'none'
+# set calib_layer
+calib_layer = None
+# set quantized_dtype
+quantized_dtype = 'auto'
+logger.info('Quantizing FP32 model Resnet18-V1')
+qsym, qarg_params, aux_params, collector = quantize_graph(sym=sym, arg_params=arg_params, aux_params=aux_params,
+ excluded_sym_names=excluded_names,
+ calib_mode=calib_mode, calib_layer=calib_layer,
+ quantized_dtype=quantized_dtype, logger=logger)
+# (optional) visualize quantized model
+mx.viz.plot_network(qsym)
+# save quantized model
+mx.model.save_checkpoint('quantized-resnet18_v1', 0, qsym, qarg_params, aux_params)
+```
+
+By applying `quantize_graph` to the symbolic model, a new quantized model can be generated, named `qsym` along with its parameters. We can see `_contrib_requantize` operators are inserted after `Convolution` to convert the INT32 output to FP32.
+
+
+
+Below table gives some descriptions.
+
+| param | type | description|
+|--------------------|-----------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
+| excluded_sym_names | list of strings | A list of strings representing the names of the symbols that users want to excluding from being quantized.|
+| calib_mode | str | If calib_mode='none', no calibration will be used and the thresholds for requantization after the corresponding layers will be calculated at runtime by calling min and max operators. The quantized models generated in this mode are normally 10-20% slower than those with calibrations during inference.
If calib_mode='naive', the min and max values of the layer outputs from a calibration dataset will be directly taken as the thresholds for quantization.
If calib_mode='entropy', the thresholds for quantization will be derived such that the KL divergence between the distributions of FP32 layer outputs and quantized layer outputs is minimized based upon the calibration dataset. |
+| calib_layer | function | Given a layer's output name in string, return True or False for deciding whether to calibrate this layer.
If yes, the statistics of the layer's output will be collected; otherwise, no information of the layer's output will be collected.
If not provided, all the layers' outputs that need requantization will be collected.|
+| quantized_dtype | str | The quantized destination type for input data. Currently support 'int8', 'uint8' and 'auto'.
'auto' means automatically select output type according to calibration result.|
+
+### Evaluate & Tune
+
+Now, you get a pair of quantized symbol and params file for inference. For Gluon inference, only difference is to load model and params by a SymbolBlock as below example:
+
+```python
+quantized_net = mx.gluon.SymbolBlock.imports('quantized-resnet18_v1-symbol.json', 'data', 'quantized-resnet18_v1-0000.params')
+quantized_net.hybridize(static_shape=True, static_alloc=True)
+batch_size = 1
+data = mx.nd.ones((batch_size,3,224,224))
+quantized_net(data)
+```
+
+Now, you can get the accuracy from a quantized network. Furthermore, you can try to select different layers or OPs to be quantized by `excluded_sym_names` parameter and figure out an acceptable accuracy.
+
+### Calibrate Model (optional for performance)
+
+The quantized model generated in previous steps can be very slow during inference since it will calculate min and max at runtime. We recommend using offline calibration for better performance by setting `calib_mode` to `naive` or `entropy`. And then calling `set_monitor_callback` api to collect layer information with a subset of the validation datasets before int8 inference.
+
+```python
+# quantization configs
+# set exclude layers
+excluded_names = []
+# set calib mode.
+calib_mode = 'naive'
+# set calib_layer
+calib_layer = None
+# set quantized_dtype
+quantized_dtype = 'auto'
+logger.info('Quantizing FP32 model resnet18-V1')
+cqsym, cqarg_params, aux_params, collector = quantize_graph(sym=sym, arg_params=arg_params, aux_params=aux_params,
+ excluded_sym_names=excluded_names,
+ calib_mode=calib_mode, calib_layer=calib_layer,
+ quantized_dtype=quantized_dtype, logger=logger)
+
+# download imagenet validation dataset
+mx.test_utils.download('http://data.mxnet.io/data/val_256_q90.rec', 'dataset.rec')
+# set rgb info for data
+mean_std = {'mean_r': 123.68, 'mean_g': 116.779, 'mean_b': 103.939, 'std_r': 58.393, 'std_g': 57.12, 'std_b': 57.375}
+# set batch size
+batch_size = 16
+# create DataIter
+data = mx.io.ImageRecordIter(path_imgrec='dataset.rec', batch_size=batch_size, data_shape=batch_shape[1:], rand_crop=False, rand_mirror=False, **mean_std)
+# create module
+mod = mx.mod.Module(symbol=sym, label_names=None, context=mx.cpu())
+mod.bind(for_training=False, data_shapes=data.provide_data, label_shapes=None)
+mod.set_params(arg_params, aux_params)
+
+# calibration configs
+# set num_calib_batches
+num_calib_batches = 5
+max_num_examples = num_calib_batches * batch_size
+# monitor FP32 Inference
+mod._exec_group.execs[0].set_monitor_callback(collector.collect, monitor_all=True)
+num_batches = 0
+num_examples = 0
+for batch in data:
+ mod.forward(data_batch=batch, is_train=False)
+ num_batches += 1
+ num_examples += batch_size
+ if num_examples >= max_num_examples:
+ break
+if logger is not None:
+ logger.info("Collected statistics from %d batches with batch_size=%d"
+ % (num_batches, batch_size))
+```
+
+After that, layer information will be filled into the `collector` returned by `quantize_graph` api. Then, you need to write the layer information into int8 model by calling `calib_graph` api.
+
+
+```python
+# write scaling factor into quantized symbol
+cqsym, cqarg_params, aux_params = calib_graph(qsym=cqsym, arg_params=arg_params, aux_params=aux_params,
+ collector=collector, calib_mode=calib_mode,
+ quantized_dtype=quantized_dtype, logger=logger)
+# (optional) visualize quantized model
+mx.viz.plot_network(cqsym)
+```
+
+Below is a quantized residual block with naive calibration. We can see `min_calib_range` and `max_calib_range` are written into `_contrib_requantize` operators.
+
+
+
+When you get a quantized model with calibration, keeping sure to call fusion api again since this can fuse some `requantize` or `dequantize` operators for further performance improvement.
+
+```python
+# perform post-quantization fusion
+cqsym = cqsym.get_backend_symbol('MKLDNN_QUANTIZE')
+# (optional) visualize post-quantized model
+mx.viz.plot_network(cqsym)
+# save quantized model
+mx.model.save_checkpoint('quantized-resnet18_v1', 0, cqsym, cqarg_params, aux_params)
+```
+
+Below is a post-quantized residual block. We can see `_contrib_requantize` operators are fused into `Convolution` operators.
+
+
+
+BTW, You can also modify the `min_calib_range` and `max_calib_range` in the JSON file directly.
+
+```
+ {
+ "op": "_sg_mkldnn_conv",
+ "name": "quantized_sg_mkldnn_conv_bn_act_6",
+ "attrs": {
+ "max_calib_range": "3.562147",
+ "min_calib_range": "0.000000",
+ "quantized": "true",
+ "with_act": "true",
+ "with_bn": "true"
+ },
+......
+```
+
+### Tips for Model Calibration
+
+#### Accuracy Tuning
+
+- Try to use `entropy` calib mode;
+
+- Try to exclude some layers which may cause obvious accuracy drop;
+
+- Change calibration dataset by setting different `num_calib_batches` or shuffle your validation dataset;
+
+#### Performance Tuning
+
+- Keep sure to perform graph fusion before quantization;
+
+- If lots of `requantize` layers exist, keep sure to perform post-quantization fusion after calibration;
+
+- Compare the MXNet profile or `MKLDNN_VERBOSE` of float32 and int8 inference;
+
+## Deploy with Python/C++
+
+MXNet also supports deploy quantized models with C++. Refer [MXNet C++ Package](https://github.com/apache/incubator-mxnet/blob/master/cpp-package/README.md) for more details.
+
+
diff --git a/example/quantization/README.md b/example/quantization/README.md
index 40f1371c33cc..61a36a415130 100644
--- a/example/quantization/README.md
+++ b/example/quantization/README.md
@@ -9,13 +9,76 @@ This folder contains examples of quantizing a FP32 model with Intel® MKL-DNN or
Model Quantization with Intel® MKL-DNN
-Intel® MKL-DNN supports quantization with subgraph features on Intel® CPU Platform and can bring performance improvements on the [Intel® Xeon® Scalable Platform](https://www.intel.com/content/www/us/en/processors/xeon/scalable/xeon-scalable-platform.html). A new quantization script `imagenet_gen_qsym_mkldnn.py` has been designed to launch quantization for CNN models with Intel® MKL-DNN. This script integrates with [Gluon-CV modelzoo](https://gluon-cv.mxnet.io/model_zoo/classification.html), so that more pre-trained models can be downloaded from Gluon-CV and then converted for quantization. This script also supports custom models.
-
-Calibration is used for generating a calibration table for the quantized symbol. The quantization script supports three methods:
-
-- **none:** No calibration will be used. The thresholds for quantization will be calculated on the fly. This will result in inference speed slowdown and loss of accuracy in general.
-- **naive:** Simply take min and max values of layer outputs as thresholds for quantization. In general, the inference accuracy worsens with more examples used in calibration. It is recommended to use `entropy` mode as it produces more accurate inference results.
-- **entropy:** Calculate KL divergence of the fp32 output and quantized output for optimal thresholds. This mode is expected to produce the best inference accuracy of all three kinds of quantized models if the calibration dataset is representative enough of the inference dataset.
+Intel® MKL-DNN supports quantization with subgraph features on Intel® CPU Platform and can bring performance improvements on the [Intel® Xeon® Scalable Platform](https://www.intel.com/content/www/us/en/processors/xeon/scalable/xeon-scalable-platform.html). A new quantization script `imagenet_gen_qsym_mkldnn.py` has been designed to launch quantization for image-classification models with Intel® MKL-DNN. This script integrates with [Gluon-CV modelzoo](https://gluon-cv.mxnet.io/model_zoo/classification.html), so that more pre-trained models can be downloaded from Gluon-CV and then converted for quantization. To apply quantization flow to your project directly, please refer [Quantize custom models with MKL-DNN backend](https://mxnet.incubator.apache.org/tutorials/mkldnn/mkldnn_quantization.html).
+
+```
+usage: imagenet_gen_qsym_mkldnn.py [-h] [--model MODEL] [--epoch EPOCH]
+ [--no-pretrained] [--batch-size BATCH_SIZE]
+ [--label-name LABEL_NAME]
+ [--calib-dataset CALIB_DATASET]
+ [--image-shape IMAGE_SHAPE]
+ [--data-nthreads DATA_NTHREADS]
+ [--num-calib-batches NUM_CALIB_BATCHES]
+ [--exclude-first-conv] [--shuffle-dataset]
+ [--shuffle-chunk-seed SHUFFLE_CHUNK_SEED]
+ [--shuffle-seed SHUFFLE_SEED]
+ [--calib-mode CALIB_MODE]
+ [--quantized-dtype {auto,int8,uint8}]
+ [--enable-calib-quantize ENABLE_CALIB_QUANTIZE]
+
+Generate a calibrated quantized model from a FP32 model with Intel MKL-DNN
+support
+
+optional arguments:
+ -h, --help show this help message and exit
+ --model MODEL model to be quantized.
+ --epoch EPOCH number of epochs, default is 0
+ --no-pretrained If enabled, will not download pretrained model from
+ MXNet or Gluon-CV modelzoo.
+ --batch-size BATCH_SIZE
+ --label-name LABEL_NAME
+ --calib-dataset CALIB_DATASET
+ path of the calibration dataset
+ --image-shape IMAGE_SHAPE
+ --data-nthreads DATA_NTHREADS
+ number of threads for data decoding
+ --num-calib-batches NUM_CALIB_BATCHES
+ number of batches for calibration
+ --exclude-first-conv excluding quantizing the first conv layer since the
+ input data may have negative value which doesn't
+ support at moment
+ --shuffle-dataset shuffle the calibration dataset
+ --shuffle-chunk-seed SHUFFLE_CHUNK_SEED
+ shuffling chunk seed, see https://mxnet.incubator.apac
+ he.org/api/python/io/io.html?highlight=imager#mxnet.io
+ .ImageRecordIter for more details
+ --shuffle-seed SHUFFLE_SEED
+ shuffling seed, see https://mxnet.incubator.apache.org
+ /api/python/io/io.html?highlight=imager#mxnet.io.Image
+ RecordIter for more details
+ --calib-mode CALIB_MODE
+ calibration mode used for generating calibration table
+ for the quantized symbol; supports 1. none: no
+ calibration will be used. The thresholds for
+ quantization will be calculated on the fly. This will
+ result in inference speed slowdown and loss of
+ accuracy in general. 2. naive: simply take min and max
+ values of layer outputs as thresholds for
+ quantization. In general, the inference accuracy
+ worsens with more examples used in calibration. It is
+ recommended to use `entropy` mode as it produces more
+ accurate inference results. 3. entropy: calculate KL
+ divergence of the fp32 output and quantized output for
+ optimal thresholds. This mode is expected to produce
+ the best inference accuracy of all three kinds of
+ quantized models if the calibration dataset is
+ representative enough of the inference dataset.
+ --quantized-dtype {auto,int8,uint8}
+ quantization destination data type for input data
+ --enable-calib-quantize ENABLE_CALIB_QUANTIZE
+ If enabled, the quantize op will be calibrated offline
+ if calibration mode is enabled
+```
Use the following command to install [Gluon-CV](https://gluon-cv.mxnet.io/):
@@ -23,12 +86,13 @@ Use the following command to install [Gluon-CV](https://gluon-cv.mxnet.io/):
pip install gluoncv
```
-The following models have been tested on Linux systems.
+Below are some quantization demos. These models have been tested on Linux systems.
| Model | Source | Dataset | FP32 Accuracy (top-1/top-5)| INT8 Accuracy (top-1/top-5)|
|:---|:---|---|:---:|:---:|
| [ResNet18-V1](#3) | [Gluon-CV](https://gluon-cv.mxnet.io/model_zoo/classification.html) | [Validation Dataset](http://data.mxnet.io/data/val_256_q90.rec) |70.15%/89.38%|69.92%/89.26%|
| [ResNet50-V1](#3) | [Gluon-CV](https://gluon-cv.mxnet.io/model_zoo/classification.html) | [Validation Dataset](http://data.mxnet.io/data/val_256_q90.rec) | 76.34%/93.13% | 75.91%/92.95% |
+| [ResNet50-V1b](#3) | [Gluon-CV](https://gluon-cv.mxnet.io/model_zoo/classification.html) | [Validation Dataset](http://data.mxnet.io/data/val_256_q90.rec) | 76.82%/93.38% | 76.39%/93.24% |
| [ResNet101-V1](#3) | [Gluon-CV](https://gluon-cv.mxnet.io/model_zoo/classification.html) | [Validation Dataset](http://data.mxnet.io/data/val_256_q90.rec) | 77.33%/93.59% | 77.05%/93.43% |
|[Squeezenet 1.0](#4)|[Gluon-CV](https://gluon-cv.mxnet.io/model_zoo/classification.html)|[Validation Dataset](http://data.mxnet.io/data/val_256_q90.rec)|56.98%/79.20%|52.98%/77.21%|
|[MobileNet 1.0](#5)|[Gluon-CV](https://gluon-cv.mxnet.io/model_zoo/classification.html)|[Validation Dataset](http://data.mxnet.io/data/val_256_q90.rec)|72.23%/90.64%|72.03%/90.42%|
@@ -39,7 +103,7 @@ The following models have been tested on Linux systems.
| [SSD-VGG16](#10) | [example/ssd](https://github.com/apache/incubator-mxnet/tree/master/example/ssd) | VOC2007/2012 | 0.8366 mAP | 0.8364 mAP |
| [SSD-VGG16](#10) | [example/ssd](https://github.com/apache/incubator-mxnet/tree/master/example/ssd) | COCO2014 | 0.2552 mAP | 0.253 mAP |
-ResNet18/50/101-V1
+ResNetV1
The following command is to download the pre-trained model from Gluon-CV and transfer it into the symbolic model which would be finally quantized. The [validation dataset](http://data.mxnet.io/data/val_256_q90.rec) is available for testing the pre-trained models:
@@ -47,7 +111,7 @@ The following command is to download the pre-trained model from Gluon-CV and tra
python imagenet_gen_qsym_mkldnn.py --model=resnet50_v1 --num-calib-batches=5 --calib-mode=naive
```
-The model would be automatically replaced in fusion and quantization format. It is then saved as the quantized symbol and parameter files in the `./model` directory. The following command is to launch inference.
+The model would be automatically replaced in fusion and quantization format. It is then saved as the quantized symbol and parameter files in the `./model` directory. Set `--model` to `resnet18_v1/resnet50_v1b/resnet101_v1` to quantize other models. The following command is to launch inference.
```
# USE MKLDNN AS SUBGRAPH BACKEND
@@ -219,17 +283,14 @@ SSD model is located in [example/ssd](https://github.com/apache/incubator-mxnet/
This script also supports custom symbolic models. You can easily add some quantization layer configs in `imagenet_gen_qsym_mkldnn.py` like below:
```
-elif args.model == 'custom':
+else:
+ logger.info('Please set proper RGB configs for model %s' % args.model)
# add rgb mean/std of your model.
rgb_mean = '0,0,0'
rgb_std = '0,0,0'
- calib_layer = lambda name: name.endswith('_output')
# add layer names you donnot want to quantize.
- # add conv/pool layer names that has negative inputs
- # since Intel® MKL-DNN only support uint8 quantization temporary.
- # add all fc layer names since Intel® MKL-DNN does not support temporary.
+ logger.info('Please set proper excluded_sym_names for model %s' % args.model)
excluded_sym_names += ['layers']
- # add your first conv layer names since Intel® MKL-DNN only support uint8 quantization temporary.
if exclude_first_conv:
excluded_sym_names += ['layers']
```
@@ -247,7 +308,7 @@ export MXNET_SUBGRAPH_BACKEND=MKLDNN
python imagenet_inference.py --symbol-file=./model/custom-symbol.json --param-file=./model/custom-0000.params --rgb-mean=* --rgb-std=* --num-skipped-batches=* --batch-size=* --num-inference-batches=*--dataset=./data/* --ctx=cpu
```
-3. Then, you should add `rgb_mean`, `rgb_std` and `excluded_sym_names` in this script. Notice that you should exclude conv/pool layers that have negative data since Intel® MKL-DNN only supports `uint8` quantization temporarily. You should also exclude all fc layers in your model.
+3. Then, you should add `rgb_mean`, `rgb_std` and `excluded_sym_names` in this script.
4. Then, you can run the following command for quantization:
diff --git a/example/quantization/imagenet_gen_qsym_mkldnn.py b/example/quantization/imagenet_gen_qsym_mkldnn.py
index 482127ba355c..302a04449885 100644
--- a/example/quantization/imagenet_gen_qsym_mkldnn.py
+++ b/example/quantization/imagenet_gen_qsym_mkldnn.py
@@ -92,21 +92,12 @@ def save_params(fname, arg_params, aux_params, logger=None):
if __name__ == '__main__':
parser = argparse.ArgumentParser(description='Generate a calibrated quantized model from a FP32 model with Intel MKL-DNN support')
- parser.add_argument('--model', type=str, choices=['resnet18_v1',
- 'resnet50_v1',
- 'resnet101_v1',
- 'inceptionv3',
- 'squeezenet1.0',
- 'mobilenet1.0',
- 'mobilenetv2_1.0',
- 'imagenet1k-resnet-152',
- 'imagenet1k-inception-bn',
- 'custom'],
- help='currently only supports imagenet1k-resnet-50_v1, imagenet1k-resnet-152 or imagenet1k-inception-bn.'
- 'you can set to custom to load your pre-trained model.')
- parser.add_argument('--use-gluon-model', type=bool, default=False,
- help='If enabled, will download pretrained model from Gluon-CV '
- 'and convert to symbolic model ')
+ parser.add_argument('--model', type=str, default='resnet50_v1',
+ help='model to be quantized.')
+ parser.add_argument('--epoch', type=int, default=0,
+ help='number of epochs, default is 0')
+ parser.add_argument('--no-pretrained', action='store_true', default=False,
+ help='If enabled, will not download pretrained model from MXNet or Gluon-CV modelzoo.')
parser.add_argument('--batch-size', type=int, default=32)
parser.add_argument('--label-name', type=str, default='softmax_label')
parser.add_argument('--calib-dataset', type=str, default='data/val_256_q90.rec',
@@ -155,6 +146,7 @@ def save_params(fname, arg_params, aux_params, logger=None):
logger = logging.getLogger('logger')
logger.setLevel(logging.INFO)
+ logger.info(args)
logger.info('shuffle_dataset=%s' % args.shuffle_dataset)
calib_mode = args.calib_mode
@@ -165,29 +157,24 @@ def save_params(fname, arg_params, aux_params, logger=None):
download_calib_dataset('http://data.mxnet.io/data/val_256_q90.rec', args.calib_dataset)
# download model
- if args.model in ['resnet18_v1',
- 'resnet50_v1',
- 'resnet101_v1',
- 'squeezenet1.0',
- 'mobilenet1.0',
- 'mobilenetv2_1.0',
- 'inceptionv3']:
- logger.info('model %s is converted from GluonCV' % args.model)
- args.use_gluon_model = True
- if args.use_gluon_model == True:
- prefix = convert_from_gluon(model_name=args.model, image_shape=args.image_shape, classes=1000, logger=logger)
- epoch = 0
- sym, arg_params, aux_params = mx.model.load_checkpoint(prefix, epoch)
- elif args.model == 'custom':
+ if not args.no_pretrained:
+ logger.info('Get pre-trained model from MXNet or Gluoncv modelzoo.')
+ logger.info('If you want to use custom model, please set --no-pretrained.')
+ if args.model in ['imagenet1k-resnet-152', 'imagenet1k-inception-bn']:
+ logger.info('model %s is downloaded from MXNet modelzoo' % args.model)
+ prefix, epoch = download_model(model_name=args.model, logger=logger)
+ else:
+ logger.info('model %s is converted from GluonCV' % args.model)
+ prefix = convert_from_gluon(model_name=args.model, image_shape=args.image_shape, classes=1000, logger=logger)
+ rgb_mean = '123.68,116.779,103.939'
+ rgb_std = '58.393, 57.12, 57.375'
+ epoch = 0
+ else:
dir_path = os.path.dirname(os.path.realpath(__file__))
prefix = os.path.join(dir_path, 'model', args.model)
- epoch = 0
- sym, arg_params, aux_params = mx.model.load_checkpoint(prefix, epoch)
- else:
- prefix, epoch = download_model(model_name=args.model, logger=logger)
- sym, arg_params, aux_params = mx.model.load_checkpoint(prefix, epoch)
+ epoch = args.epoch
- sym = sym.get_backend_symbol('MKLDNN_QUANTIZE')
+ sym, arg_params, aux_params = mx.model.load_checkpoint(prefix, epoch)
# get batch size
batch_size = args.batch_size
@@ -212,57 +199,59 @@ def save_params(fname, arg_params, aux_params, logger=None):
logger.info('quantized dtype is set to uint8, will exclude first conv.')
exclude_first_conv = True
excluded_sym_names = []
- if args.model == 'imagenet1k-resnet-152':
- rgb_mean = '0,0,0'
- rgb_std = '1,1,1'
- excluded_sym_names += ['flatten0']
- if exclude_first_conv:
- excluded_sym_names += ['conv0']
- elif args.model == 'imagenet1k-inception-bn':
- rgb_mean = '123.68,116.779,103.939'
- rgb_std = '1,1,1'
- excluded_sym_names += ['flatten']
- if exclude_first_conv:
- excluded_sym_names += ['conv_1']
- elif args.model in ['resnet18_v1', 'resnet50_v1', 'resnet101_v1']:
- rgb_mean = '123.68,116.779,103.939'
- rgb_std = '58.393, 57.12, 57.375'
- if exclude_first_conv:
- excluded_sym_names += ['resnetv10_conv0_fwd']
- elif args.model == 'squeezenet1.0':
- rgb_mean = '123.68,116.779,103.939'
- rgb_std = '58.393, 57.12, 57.375'
- excluded_sym_names += ['squeezenet0_flatten0_flatten0']
- if exclude_first_conv:
- excluded_sym_names += ['squeezenet0_conv0_fwd']
- elif args.model == 'mobilenet1.0':
- rgb_mean = '123.68,116.779,103.939'
- rgb_std = '58.393, 57.12, 57.375'
- excluded_sym_names += ['mobilenet0_flatten0_flatten0',
- 'mobilenet0_pool0_fwd']
- if exclude_first_conv:
- excluded_sym_names += ['mobilenet0_conv0_fwd']
- elif args.model == 'mobilenetv2_1.0':
- rgb_mean = '123.68,116.779,103.939'
- rgb_std = '58.393, 57.12, 57.375'
- excluded_sym_names += ['mobilenetv20_output_flatten0_flatten0']
- if exclude_first_conv:
- excluded_sym_names += ['mobilenetv20_conv0_fwd']
- elif args.model == 'inceptionv3':
- rgb_mean = '123.68,116.779,103.939'
- rgb_std = '58.393, 57.12, 57.375'
- if exclude_first_conv:
- excluded_sym_names += ['inception30_conv0_fwd']
- elif args.model == 'custom':
+ if not args.no_pretrained:
+ if args.model == 'imagenet1k-resnet-152':
+ rgb_mean = '0,0,0'
+ rgb_std = '1,1,1'
+ excluded_sym_names += ['flatten0']
+ if exclude_first_conv:
+ excluded_sym_names += ['conv0']
+ elif args.model == 'imagenet1k-inception-bn':
+ rgb_mean = '123.68,116.779,103.939'
+ rgb_std = '1,1,1'
+ excluded_sym_names += ['flatten']
+ if exclude_first_conv:
+ excluded_sym_names += ['conv_1']
+ elif args.model.find('resnet') != -1 and args.model.find('v1') != -1:
+ if exclude_first_conv:
+ excluded_sym_names += ['resnetv10_conv0_fwd']
+ elif args.model.find('resnet') != -1 and args.model.find('v2') != -1:
+ excluded_sym_names += ['resnetv20_flatten0_flatten0']
+ if exclude_first_conv:
+ excluded_sym_names += ['resnetv20_conv0_fwd']
+ elif args.model.find('vgg') != -1:
+ if exclude_first_conv:
+ excluded_sym_names += ['vgg0_conv0_fwd']
+ elif args.model.find('squeezenet1') != -1:
+ excluded_sym_names += ['squeezenet0_flatten0_flatten0']
+ if exclude_first_conv:
+ excluded_sym_names += ['squeezenet0_conv0_fwd']
+ elif args.model.find('mobilenet') != -1 and args.model.find('v2') == -1:
+ excluded_sym_names += ['mobilenet0_flatten0_flatten0',
+ 'mobilenet0_pool0_fwd']
+ if exclude_first_conv:
+ excluded_sym_names += ['mobilenet0_conv0_fwd']
+ elif args.model.find('mobilenet') != -1 and args.model.find('v2') != -1:
+ excluded_sym_names += ['mobilenetv20_output_flatten0_flatten0']
+ if exclude_first_conv:
+ excluded_sym_names += ['mobilenetv20_conv0_fwd']
+ elif args.model == 'inceptionv3':
+ if exclude_first_conv:
+ excluded_sym_names += ['inception30_conv0_fwd']
+ else:
+ raise ValueError('Currently, model %s is not supported in this script' % args.model)
+ else:
+ logger.info('Please set proper RGB configs for model %s' % args.model)
# add rgb mean/std of your model.
rgb_mean = '0,0,0'
rgb_std = '0,0,0'
# add layer names you donnot want to quantize.
+ logger.info('Please set proper excluded_sym_names for model %s' % args.model)
excluded_sym_names += ['layers']
if exclude_first_conv:
excluded_sym_names += ['layers']
- else:
- raise ValueError('model %s is not supported in this script' % args.model)
+
+ logger.info('These layers have been excluded %s' % excluded_sym_names)
label_name = args.label_name
logger.info('label_name = %s' % label_name)
@@ -281,10 +270,10 @@ def save_params(fname, arg_params, aux_params, logger=None):
combine_mean_std.update(std_args)
if calib_mode == 'none':
logger.info('Quantizing FP32 model %s' % args.model)
- qsym, qarg_params, aux_params = quantize_model(sym=sym, arg_params=arg_params, aux_params=aux_params,
- ctx=ctx, excluded_sym_names=excluded_sym_names,
- calib_mode=calib_mode, quantized_dtype=args.quantized_dtype,
- logger=logger)
+ qsym, qarg_params, aux_params = quantize_model_mkldnn(sym=sym, arg_params=arg_params, aux_params=aux_params,
+ ctx=ctx, excluded_sym_names=excluded_sym_names,
+ calib_mode=calib_mode, quantized_dtype=args.quantized_dtype,
+ logger=logger)
sym_name = '%s-symbol.json' % (prefix + '-quantized')
else:
logger.info('Creating ImageRecordIter for reading calibration dataset')
@@ -301,12 +290,12 @@ def save_params(fname, arg_params, aux_params, logger=None):
seed=args.shuffle_seed,
**combine_mean_std)
- qsym, qarg_params, aux_params = quantize_model(sym=sym, arg_params=arg_params, aux_params=aux_params,
- ctx=ctx, excluded_sym_names=excluded_sym_names,
- calib_mode=calib_mode, calib_data=data,
- num_calib_examples=num_calib_batches * batch_size,
- calib_layer=calib_layer, quantized_dtype=args.quantized_dtype,
- label_names=(label_name,), logger=logger)
+ qsym, qarg_params, aux_params = quantize_model_mkldnn(sym=sym, arg_params=arg_params, aux_params=aux_params,
+ ctx=ctx, excluded_sym_names=excluded_sym_names,
+ calib_mode=calib_mode, calib_data=data,
+ num_calib_examples=num_calib_batches * batch_size,
+ calib_layer=calib_layer, quantized_dtype=args.quantized_dtype,
+ label_names=(label_name,), logger=logger)
if calib_mode == 'entropy':
suffix = '-quantized-%dbatches-entropy' % num_calib_batches
elif calib_mode == 'naive':
@@ -315,7 +304,6 @@ def save_params(fname, arg_params, aux_params, logger=None):
raise ValueError('unknow calibration mode %s received, only supports `none`, `naive`, and `entropy`'
% calib_mode)
sym_name = '%s-symbol.json' % (prefix + suffix)
- qsym = qsym.get_backend_symbol('MKLDNN_QUANTIZE')
save_symbol(sym_name, qsym, logger)
param_name = '%s-%04d.params' % (prefix + '-quantized', epoch)
save_params(param_name, qarg_params, aux_params, logger)
diff --git a/python/mxnet/contrib/quantization.py b/python/mxnet/contrib/quantization.py
index b94b5a8da32a..fa2ab1842f5f 100644
--- a/python/mxnet/contrib/quantization.py
+++ b/python/mxnet/contrib/quantization.py
@@ -543,3 +543,240 @@ def quantize_model(sym, arg_params, aux_params,
qarg_params = _quantize_params(qsym, arg_params, th_dict)
return qsym, qarg_params, aux_params
+
+def quantize_model_mkldnn(sym, arg_params, aux_params,
+ data_names=('data',), label_names=('softmax_label',),
+ ctx=cpu(), excluded_sym_names=None, calib_mode='entropy',
+ calib_data=None, num_calib_examples=None, calib_layer=None,
+ quantized_dtype='int8', logger=logging):
+ """User-level API for generating a fusion + quantized model from a FP32 model
+ w/ or w/o calibration with Intel MKL-DNN.
+ The backend quantized operators are only enabled for Linux systems. Please do not run
+ inference using the quantized models on Windows for now.
+
+ Parameters
+ ----------
+ sym : str or Symbol
+ Defines the structure of a neural network for FP32 data types.
+ arg_params : dict
+ Dictionary of name to `NDArray`.
+ aux_params : dict
+ Dictionary of name to `NDArray`.
+ data_names : a list of strs
+ Data names required for creating a Module object to run forward propagation on the
+ calibration dataset.
+ label_names : a list of strs
+ Label names required for creating a Module object to run forward propagation on the
+ calibration dataset.
+ ctx : Context
+ Defines the device that users want to run forward propagation on the calibration
+ dataset for collecting layer output statistics. Currently, only supports single context.
+ excluded_sym_names : list of strings
+ A list of strings representing the names of the symbols that users want to excluding
+ from being quantized.
+ calib_mode : str
+ If calib_mode='none', no calibration will be used and the thresholds for
+ requantization after the corresponding layers will be calculated at runtime by
+ calling min and max operators. The quantized models generated in this
+ mode are normally 10-20% slower than those with calibrations during inference.
+ If calib_mode='naive', the min and max values of the layer outputs from a calibration
+ dataset will be directly taken as the thresholds for quantization.
+ If calib_mode='entropy' (default mode), the thresholds for quantization will be
+ derived such that the KL divergence between the distributions of FP32 layer outputs and
+ quantized layer outputs is minimized based upon the calibration dataset.
+ calib_data : DataIter
+ A data iterator initialized by the calibration dataset.
+ num_calib_examples : int or None
+ The maximum number of examples that user would like to use for calibration. If not provided,
+ the whole calibration dataset will be used.
+ calib_layer : function
+ Given a layer's output name in string, return True or False for deciding whether to
+ calibrate this layer. If yes, the statistics of the layer's output will be collected;
+ otherwise, no information of the layer's output will be collected. If not provided,
+ all the layers' outputs that need requantization will be collected.
+ quantized_dtype : str
+ The quantized destination type for input data. Currently support 'int8'
+ , 'uint8' and 'auto'. 'auto' means automatically select output type according to calibration result.
+ Default value is 'int8'.
+ logger : Object
+ A logging object for printing information during the process of quantization.
+
+ Returns
+ -------
+ tuple
+ A tuple of quantized symbol, quantized arg_params, and aux_params.
+ -------
+ """
+ if ctx != cpu():
+ raise ValueError(
+ 'quantize_model_mkldnn only support Intel cpu platform with MKL-DNN Backend')
+
+ sym = sym.get_backend_symbol('MKLDNN_QUANTIZE')
+
+ qsym, qarg_params, aux_params = quantize_model(sym=sym, arg_params=arg_params, aux_params=aux_params,
+ data_names=data_names, label_names=label_names,
+ ctx=ctx, excluded_sym_names=excluded_sym_names,
+ calib_mode=calib_mode, calib_data=calib_data,
+ num_calib_examples=num_calib_examples, calib_layer=calib_layer,
+ quantized_dtype=quantized_dtype, logger=logger)
+
+ qsym = qsym.get_backend_symbol('MKLDNN_QUANTIZE')
+
+ return qsym, qarg_params, aux_params
+
+def quantize_graph(sym, arg_params, aux_params,
+ excluded_sym_names=None, calib_mode='entropy',
+ calib_layer=None, quantized_dtype='int8', logger=logging):
+ """User-level API for generating a quantized model from a FP32 model w/o calibration
+ and a collector for naive or entropy calibration.
+ The backend quantized operators are only enabled for Linux systems. Please do not run
+ inference using the quantized models on Windows for now.
+ The quantization implementation adopts the TensorFlow's approach:
+ https://www.tensorflow.org/performance/quantization.
+ The calibration implementation borrows the idea of Nvidia's 8-bit Inference with TensorRT:
+ http://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inference-with-tensorrt.pdf
+ and adapts the method to MXNet.
+ Parameters
+ ----------
+ sym : str or Symbol
+ Defines the structure of a neural network for FP32 data types.
+ arg_params : dict
+ Dictionary of name to `NDArray`.
+ aux_params : dict
+ Dictionary of name to `NDArray`.
+ excluded_sym_names : list of strings
+ A list of strings representing the names of the symbols that users want to excluding
+ from being quantized.
+ calib_mode : str
+ If calib_mode='none', no calibration will be used and the thresholds for
+ requantization after the corresponding layers will be calculated at runtime by
+ calling min and max operators. The quantized models generated in this
+ mode are normally 10-20% slower than those with calibrations during inference.
+ If calib_mode='naive', the min and max values of the layer outputs from a calibration
+ dataset will be directly taken as the thresholds for quantization.
+ If calib_mode='entropy' (default mode), the thresholds for quantization will be
+ derived such that the KL divergence between the distributions of FP32 layer outputs and
+ quantized layer outputs is minimized based upon the calibration dataset.
+ calib_layer : function
+ Given a layer's output name in string, return True or False for deciding whether to
+ calibrate this layer. If yes, the statistics of the layer's output will be collected;
+ otherwise, no information of the layer's output will be collected. If not provided,
+ all the layers' outputs that need requantization will be collected.
+ quantized_dtype : str
+ The quantized destination type for input data. Currently support 'int8'
+ , 'uint8' and 'auto'. 'auto' means automatically select output type according to calibration result.
+ Default value is 'int8'.
+ logger : Object
+ A logging object for printing information during the process of quantization.
+ Returns
+ -------
+ tuple
+ A tuple of quantized symbol, quantized arg_params, aux_params and collector.
+ -------
+ """
+ if excluded_sym_names is None:
+ excluded_sym_names = []
+ if not isinstance(excluded_sym_names, list):
+ raise ValueError('excluded_sym_names must be a list of strings representing'
+ ' the names of the symbols that will not be quantized,'
+ ' while received type %s' % str(type(excluded_sym_names)))
+
+ logger.info('Quantizing graph')
+ if quantized_dtype not in ('int8', 'uint8', 'auto'):
+ raise ValueError('unknown quantized_dtype %s received,'
+ ' expected `int8`, `uint8` or `auto`' % quantized_dtype)
+ qsym = _quantize_symbol(sym, excluded_symbols=excluded_sym_names,
+ offline_params=list(arg_params.keys()),
+ quantized_dtype=quantized_dtype)
+
+ th_dict = {}
+ collector = None
+ if calib_mode is not None and calib_mode != 'none':
+ if calib_mode == 'entropy':
+ collector = _LayerOutputCollector(
+ include_layer=calib_layer, logger=logger)
+ logger.info(
+ 'Create a layer output collector for entropy calibration.')
+ elif calib_mode == 'naive':
+ collector = _LayerOutputMinMaxCollector(
+ include_layer=calib_layer, logger=logger)
+ logger.info(
+ 'Create a layer output minmax collector for naive calibration')
+ else:
+ raise ValueError('unknown calibration mode %s received,'
+ ' expected `none`, `naive`, or `entropy`' % calib_mode)
+ logger.info('Collector created, please use set_monitor_callback'
+ ' to collect calibration information.')
+
+ logger.info('Quantizing parameters')
+ qarg_params = _quantize_params(qsym, arg_params, th_dict)
+
+ return qsym, qarg_params, aux_params, collector
+
+def calib_graph(qsym, arg_params, aux_params, collector,
+ calib_mode='entropy', quantized_dtype='int8', logger=logging):
+ """User-level API for calibrating a quantized model using a filled collector.
+ The backend quantized operators are only enabled for Linux systems. Please do not run
+ inference using the quantized models on Windows for now.
+ The quantization implementation adopts the TensorFlow's approach:
+ https://www.tensorflow.org/performance/quantization.
+ The calibration implementation borrows the idea of Nvidia's 8-bit Inference with TensorRT:
+ http://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inference-with-tensorrt.pdf
+ and adapts the method to MXNet.
+ Parameters
+ ----------
+ qsym : str or Symbol
+ Defines the structure of a neural network for INT8 data types.
+ arg_params : dict
+ Dictionary of name to `NDArray`.
+ aux_params : dict
+ Dictionary of name to `NDArray`.
+ collector : function
+ layer collector for naive or entropy calibration.
+ calib_mode : str
+ If calib_mode='none', no calibration will be used and the thresholds for
+ requantization after the corresponding layers will be calculated at runtime by
+ calling min and max operators. The quantized models generated in this
+ mode are normally 10-20% slower than those with calibrations during inference.
+ If calib_mode='naive', the min and max values of the layer outputs from a calibration
+ dataset will be directly taken as the thresholds for quantization.
+ If calib_mode='entropy' (default mode), the thresholds for quantization will be
+ derived such that the KL divergence between the distributions of FP32 layer outputs and
+ quantized layer outputs is minimized based upon the calibration dataset.
+ calib_layer : function
+ Given a layer's output name in string, return True or False for deciding whether to
+ calibrate this layer. If yes, the statistics of the layer's output will be collected;
+ otherwise, no information of the layer's output will be collected. If not provided,
+ all the layers' outputs that need requantization will be collected.
+ quantized_dtype : str
+ The quantized destination type for input data. Currently support 'int8'
+ , 'uint8' and 'auto'. 'auto' means automatically select output type according to calibration result.
+ Default value is 'int8'.
+ logger : Object
+ A logging object for printing information during the process of quantization.
+ Returns
+ -------
+ tuple
+ A tuple of calibrated symbol, quantized arg_params, aux_params.
+ -------
+ """
+ th_dict = {}
+ if calib_mode is not None and calib_mode != 'none':
+ if calib_mode == 'entropy':
+ logger.info('Calculating optimal thresholds for quantization')
+ th_dict = _get_optimal_thresholds(
+ collector.nd_dict, quantized_dtype, logger=logger)
+ elif calib_mode == 'naive':
+ th_dict = collector.min_max_dict
+ else:
+ raise ValueError('unknown calibration mode %s received,'
+ ' expected `none`, `naive`, or `entropy`' % calib_mode)
+ logger.info('Calibrating quantized symbol')
+ qsym = _calibrate_quantized_sym(qsym, th_dict)
+ else:
+ raise ValueError('please set calibration mode to naive or entropy.')
+
+ logger.info('Quantizing parameters')
+ qarg_params = _quantize_params(qsym, arg_params, th_dict)
+
+ return qsym, qarg_params, aux_params
diff --git a/tests/tutorials/test_tutorials.py b/tests/tutorials/test_tutorials.py
index b5f84d550636..21e0f274d219 100644
--- a/tests/tutorials/test_tutorials.py
+++ b/tests/tutorials/test_tutorials.py
@@ -207,3 +207,6 @@ def test_control_flow():
def test_amp():
assert _test_tutorial_nb('amp/amp_tutorial')
+
+def test_mkldnn_quantization():
+ assert _test_tutorial_nb('mkldnn/mkldnn_quantization')
\ No newline at end of file