AION
Research paperEfficiency & Inference1 source · Oct 7, 2026

ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals

Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count.

Key points

  • KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes.
  • Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals.
  • Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction.
  • Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization.

Sources (1)

  • [1]ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:40 PM
    Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count.
    KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes.

Extractive summary: sentences quoted from the sources.