ResearchResearch paperLarge Language Models · Interpretability1 source · Oct 7, 2026

Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values

Based on this view, we propose Matrix Approximation Sparse Attention (MASA).

Key points

  • Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length.
  • Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix.
  • We argue that this is the core conceptual issue: sparse attention should be formulated as matrix approximation, not as blindly choosing the largest values from a bag of entries.
  • Extensive experiments across multiple sparse attention methods, benchmarks, and LLM backbones show consistent accuracy gains, supporting both MASA and the matrix-approximation view of sparse attention.

Sources (1)

  • [1]Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:17 PM
    Based on this view, we propose Matrix Approximation Sparse Attention (MASA).
    Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length.

Extractive summary: sentences quoted from the sources.

Related