AION
Research paperEfficiency & Inference1 source · Oct 8, 2026

Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration

KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference.

Key points

  • KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers.
  • This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention.
  • We identify RAM-Net as such a bridge through soft assignments over a discrete address space.
  • Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation.

Sources (1)

  • [1]Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:12 AM
    KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference.
    KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers.

Extractive summary: sentences quoted from the sources.