From 6f14fea5721a68b4ae21c90bac9fc70456509015 Mon Sep 17 00:00:00 2001 From: yi111 <153097222+Yi-111-a@users.noreply.github.com> Date: Fri, 25 Sep 2026 04:17:07 +0800 Subject: [PATCH 1/3] docs(onnx): clarify INT4 calibration data guidance Signed-off-by: yi111 <153097222+Yi-111-a@users.noreply.github.com> --- examples/onnx_ptq/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/examples/onnx_ptq/README.md b/examples/onnx_ptq/README.md index 7ee260d8765..f59cf6e2e18 100644 --- a/examples/onnx_ptq/README.md +++ b/examples/onnx_ptq/README.md @@ -103,7 +103,7 @@ python image_prep.py \ --fp16 # ``` -> *For Int4 quantization, it is recommended to set `--calibration_data_size=64`.* +> *There is no INT4-specific image-count requirement. The `awq_clip` method uses calibration data to search weight-clipping parameters, while `rtn_dq` does not use calibration data. Choose a representative dataset and validate the quantized model's accuracy for your model and resource constraints.* ### Quantize ONNX Model to FP8, INT8 or INT4 From 39462c10677240fd6c28e5287227c18974b79356 Mon Sep 17 00:00:00 2001 From: Yi-111-a <153097222+Yi-111-a@users.noreply.github.com> Date: Mon, 28 Sep 2026 00:25:18 +0800 Subject: [PATCH 2/3] docs(onnx): scope calibration-data guidance to methods that use it rtn_dq ignores calibration data, so say so where the calibration-data step is introduced and note that the data path can be omitted for it. Name all AWQ methods as the INT4 consumers of calibration data. Co-Authored-By: Claude Opus 5.5 Signed-off-by: Yi-111-a <153097222+Yi-111-a@users.noreply.github.com> --- examples/onnx_ptq/README.md | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/examples/onnx_ptq/README.md b/examples/onnx_ptq/README.md index f59cf6e2e18..6d79702f6ef 100644 --- a/examples/onnx_ptq/README.md +++ b/examples/onnx_ptq/README.md @@ -92,7 +92,7 @@ python download_example_onnx.py \ ### Prepare calibration data -Calibration data is a representative subset of your training or validation dataset used during quantization to determine the optimal scale factors for converting floating-point values to lower precision formats (INT8, FP8, INT4). This data helps maintain model accuracy after quantization by analyzing the distribution of activations throughout the network. +Calibration data is a representative subset of your training or validation dataset used during quantization. For FP8 and INT8, and for the INT4 AWQ methods, it determines the optimal scale factors for converting floating-point values to lower precision formats. This data helps maintain model accuracy after quantization by analyzing the distribution of activations throughout the network. INT4 `rtn_dq` does not use calibration data, so you can skip this step for it. First, prepare some calibration data. TensorRT recommends calibration data size to be at least 500 for CNN and ViT models. The following command picks up 500 images from the [tiny-imagenet](https://huggingface.co/datasets/zh-plus/tiny-imagenet) dataset and converts them to a numpy-format calibration array. Reduce the calibration data size for resource constrained environments. @@ -103,7 +103,7 @@ python image_prep.py \ --fp16 # ``` -> *There is no INT4-specific image-count requirement. The `awq_clip` method uses calibration data to search weight-clipping parameters, while `rtn_dq` does not use calibration data. Choose a representative dataset and validate the quantized model's accuracy for your model and resource constraints.* +> *There is no INT4-specific image-count requirement. The AWQ methods (`awq_clip`, `awq_lite`, `awq_full`) use calibration data to search scaling and weight-clipping parameters, while `rtn_dq` does not use calibration data. For AWQ, choose a representative dataset and validate the quantized model's accuracy for your model and resource constraints.* ### Quantize ONNX Model to FP8, INT8 or INT4 @@ -124,6 +124,8 @@ python -m modelopt.onnx.quantization \ --output_path=vit_base_patch16_224.quant.onnx ``` +> *With `--calibration_method=rtn_dq`, `--calibration_data_path` (or `calibration_data` in the Python API) can be omitted; the calibration data is not used.* + #### Option 2: Python API ```python From 211bbe3cde762db679dd53b2122b4666be121821 Mon Sep 17 00:00:00 2001 From: Yi-111-a <153097222+Yi-111-a@users.noreply.github.com> Date: Mon, 28 Sep 2026 15:06:15 +0800 Subject: [PATCH 3/3] docs(onnx): describe what each INT4 AWQ method searches Co-Authored-By: Claude Opus 5.5 Signed-off-by: Yi-111-a <153097222+Yi-111-a@users.noreply.github.com> --- examples/onnx_ptq/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/examples/onnx_ptq/README.md b/examples/onnx_ptq/README.md index 6d79702f6ef..282fa268543 100644 --- a/examples/onnx_ptq/README.md +++ b/examples/onnx_ptq/README.md @@ -103,7 +103,7 @@ python image_prep.py \ --fp16 # ``` -> *There is no INT4-specific image-count requirement. The AWQ methods (`awq_clip`, `awq_lite`, `awq_full`) use calibration data to search scaling and weight-clipping parameters, while `rtn_dq` does not use calibration data. For AWQ, choose a representative dataset and validate the quantized model's accuracy for your model and resource constraints.* +> *There is no INT4-specific image-count requirement. The AWQ methods use calibration data: `awq_clip` searches weight-clipping parameters, `awq_lite` searches scaling parameters, and `awq_full` searches scaling with weight clipping. `rtn_dq` does not use calibration data. For AWQ, choose a representative dataset and validate the quantized model's accuracy for your model and resource constraints.* ### Quantize ONNX Model to FP8, INT8 or INT4