AION
Research paperEfficiency & Inference1 source · Oct 6, 2026

A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention

Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference.

Key points

  • The large KV-cache size of modern LLMs creates a barrier to efficient deployment.
  • In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention.
  • The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, $adaptive$ pruning criterion, removing tokens that contribute least to attention computation.
  • Experimentally, Universal Attention achieves state-of-the-art $10\times$ compression on natural language and synthetic task data, while $improving$ downstream performance compared to both state-of-the-art baselines and unpruned oracles.

Sources (1)

  • [1]A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:00 PM
    Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference.
    The large KV-cache size of modern LLMs creates a barrier to efficient deployment.

Extractive summary: sentences quoted from the sources.