AION
Research paperLarge Language Models1 source · Oct 6, 2026

Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models

Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training.

Key points

  • We show that existing compensators are limited by two shared simplifications.
  • We propose a two-stage closed-form framework that removes both simplifications.
  • Stage 2 re-measures statistics on the compensated model and applies a rank-constrained natural-gradient step that absorbs the remaining first-order signal.
  • At 2 bits under QuIP#, our method reduces WikiText-2 perplexity from 12.43 to 10.26 on Qwen3-8B and from 21.11 to 13.22 on Qwen3-4B.

Sources (1)

Extractive summary: sentences quoted from the sources.