Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models
Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training.
Key points
- We show that existing compensators are limited by two shared simplifications.
- We propose a two-stage closed-form framework that removes both simplifications.
- Stage 2 re-measures statistics on the compensated model and applies a rank-constrained natural-gradient step that absorbs the remaining first-order signal.
- At 2 bits under QuIP#, our method reduces WikiText-2 perplexity from 12.43 to 10.26 on Qwen3-8B and from 21.11 to 13.22 on Qwen3-4B.
Sources (1)
- [1]Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:16 AM
Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training.
We show that existing compensators are limited by two shared simplifications.
Extractive summary: sentences quoted from the sources.