AION
Research paperEfficiency & Inference1 source · Oct 8, 2026

When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization

Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions.

Key points

  • Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision.
  • We show that the weights favored by minimizing this loss need not yield better model performance on new tasks.
  • In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data.
  • These results establish DRQ as a general post-hoc refinement framework for weight-only PTQ, achieving better downstream performance without adding inference overhead.

Sources (1)

  • [1]When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:22 AM
    Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions.
    Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision.

Extractive summary: sentences quoted from the sources.