AION
Research paperEfficiency & Inference1 source · Oct 8, 2026

VFold: Symmetry-Aware Cross-Layer Value Cache Compression

In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding.

Key points

  • While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths.
  • One solution is to compress this memory by exploiting inter-layer cache similarities.
  • Furthermore, we show that this approach can be exploited alongside existing cache compression techniques, composing with high-ratio quantization or key cache pruning to reach compression ratios that neither method reaches alone, with minimal additional cost.
  • Ultimately, our findings reveal a major source of underutilized capacity in the value cache, offering a simple yet highly effective direction for scaling context windows under memory constraints.

Sources (1)

  • [1]VFold: Symmetry-Aware Cross-Layer Value Cache Compression
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:14 PM
    In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding.
    While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths.

Extractive summary: sentences quoted from the sources.