AION
Research paperEfficiency & Inference1 source · Oct 7, 2026

Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization

Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set.

Key points

  • Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted.
  • To design such a method, we derive a layerwise probabilistic analysis of the quantization error that separates propagated error from the local perturbation introduced at a given layer.
  • On denoising tasks with DRUNet, with an average budget of 4 bits per weight, our method matches or improves state-of-the-art mixed-precision baselines under clean calibration, and is more robust to corrupted calibration, with PSNR gains of up to 7.5 dB under the tested corruptions.
  • For quantized diffusion models, our experiments show that a direct application of our framework also improves the state-of-the-art.

Sources (1)

  • [1]Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:37 AM
    Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set.
    Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted.

Extractive summary: sentences quoted from the sources.