Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization
Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set.
Key points
- Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted.
- To design such a method, we derive a layerwise probabilistic analysis of the quantization error that separates propagated error from the local perturbation introduced at a given layer.
- On denoising tasks with DRUNet, with an average budget of 4 bits per weight, our method matches or improves state-of-the-art mixed-precision baselines under clean calibration, and is more robust to corrupted calibration, with PSNR gains of up to 7.5 dB under the tested corruptions.
- For quantized diffusion models, our experiments show that a direct application of our framework also improves the state-of-the-art.
Sources (1)
- [1]Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training QuantizationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:37 AM
Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set.
Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted.
Extractive summary: sentences quoted from the sources.