A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention
Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference.
Key points
- The large KV-cache size of modern LLMs creates a barrier to efficient deployment.
- In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention.
- The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, $adaptive$ pruning criterion, removing tokens that contribute least to attention computation.
- Experimentally, Universal Attention achieves state-of-the-art $10\times$ compression on natural language and synthetic task data, while $improving$ downstream performance compared to both state-of-the-art baselines and unpruned oracles.
Sources (1)
- [1]A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal AttentionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:00 PM
Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference.
The large KV-cache size of modern LLMs creates a barrier to efficient deployment.
Extractive summary: sentences quoted from the sources.