Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions examples/onnx_ptq/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,7 @@ python download_example_onnx.py \

### Prepare calibration data

Calibration data is a representative subset of your training or validation dataset used during quantization to determine the optimal scale factors for converting floating-point values to lower precision formats (INT8, FP8, INT4). This data helps maintain model accuracy after quantization by analyzing the distribution of activations throughout the network.
Calibration data is a representative subset of your training or validation dataset used during quantization. For FP8 and INT8, and for the INT4 AWQ methods, it determines the optimal scale factors for converting floating-point values to lower precision formats. This data helps maintain model accuracy after quantization by analyzing the distribution of activations throughout the network. INT4 `rtn_dq` does not use calibration data, so you can skip this step for it.

First, prepare some calibration data. TensorRT recommends calibration data size to be at least 500 for CNN and ViT models. The following command picks up 500 images from the [tiny-imagenet](https://huggingface.co/datasets/zh-plus/tiny-imagenet) dataset and converts them to a numpy-format calibration array. Reduce the calibration data size for resource constrained environments.

Expand All @@ -103,7 +103,7 @@ python image_prep.py \
--fp16 # <Optional, if the input ONNX is in FP16 precision>
```

> *For Int4 quantization, it is recommended to set `--calibration_data_size=64`.*
> *There is no INT4-specific image-count requirement. The AWQ methods use calibration data: `awq_clip` searches weight-clipping parameters, `awq_lite` searches scaling parameters, and `awq_full` searches scaling with weight clipping. `rtn_dq` does not use calibration data. For AWQ, choose a representative dataset and validate the quantized model's accuracy for your model and resource constraints.*

### Quantize ONNX Model to FP8, INT8 or INT4

Comment thread
coderabbitai[bot] marked this conversation as resolved.
Expand All @@ -124,6 +124,8 @@ python -m modelopt.onnx.quantization \
--output_path=vit_base_patch16_224.quant.onnx
```

> *With `--calibration_method=rtn_dq`, `--calibration_data_path` (or `calibration_data` in the Python API) can be omitted; the calibration data is not used.*

#### Option 2: Python API

```python
Expand Down