VFold: Symmetry-Aware Cross-Layer Value Cache Compression
In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding.
Key points
- While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths.
- One solution is to compress this memory by exploiting inter-layer cache similarities.
- Furthermore, we show that this approach can be exploited alongside existing cache compression techniques, composing with high-ratio quantization or key cache pruning to reach compression ratios that neither method reaches alone, with minimal additional cost.
- Ultimately, our findings reveal a major source of underutilized capacity in the value cache, offering a simple yet highly effective direction for scaling context windows under memory constraints.
Sources (1)
- [1]VFold: Symmetry-Aware Cross-Layer Value Cache CompressionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:14 PM
In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding.
While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths.
Extractive summary: sentences quoted from the sources.