resolve comments

willg-nv · willg-nv · commit 8e81fe0acabf · 2026-03-09T05:58:56.000Z
Signed-off-by: Will Guo &lt;willg@nvidia.com&gt;
diff --git a/CHANGELOG.rst b/CHANGELOG.rst
@@ -20,6 +20,7 @@ NVIDIA Model Optimizer Changelog
 - Add ``nvfp4_omlp_only`` quantization format for NVFP4 quantization. This is similar to ``nvfp4_mlp_only`` but also quantizes the output projection layer in attention.
 - ``pass_through_bwd`` in the quantization config is now default to True. Please set it to False if you want to use STE with zeroed outlier gradients for potentially better QAT accuracy.
 - Add :meth:`compute_quantization_mse <modelopt.torch.quantization.model_quant.compute_quantization_mse>` API to measure per-quantizer mean-squared quantization error, with flexible wildcard and callable filtering.
+- **AutoQDQ**: New tool for automated Q/DQ (Quantize/Dequantize) placement optimization for ONNX models. Uses TensorRT latency measurements to choose insertion schemes that minimize inference time. Discovers regions automatically, groups them by structural pattern, and tests multiple Q/DQ schemes per pattern. Supports INT8 and FP8 quantization, pattern cache for warm-start on similar models, checkpoint/resume, and importing patterns from an existing QDQ baseline. CLI: ``python -m modelopt.onnx.quantization.autotune``. See the AutoQDQ guide in the documentation.
 
 **Misc**
 
diff --git a/pyproject.toml b/pyproject.toml
@@ -48,7 +48,6 @@ dependencies = [
 [project.optional-dependencies]
 onnx = [
     "cppimport",
-    "cuda-python",
     "cupy-cuda12x; platform_machine != 'aarch64' and platform_system != 'Darwin'",
     "lief",
     "ml_dtypes",