AION
Research paperEfficiency & Inference1 source · Oct 8, 2026

PageWeaver: KV-Guided Query Unions for Sparse Attention

Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work.

Key points

  • Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together.
  • We present PageWeaver, an execution design that uses selected-page affinity to assemble query groups while preserving each query's original support and complete output ownership.
  • A bounded GPU search produces query IDs, and an ID-aware two-CTA kernel consumes them without materializing reordered Q tensors or cross-page partial outputs.
  • A direct KV-page union implementation provides a complementary design study of nonlocal reuse and reduction cost.

Sources (1)

  • [1]PageWeaver: KV-Guided Query Unions for Sparse Attention
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:04 AM
    Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work.
    Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026The Missing Fourth Term for the Emulation Tensor Memory Equilibrium (TME) Model: The Residue Deconstruction Cost
  2. Oct 7, 2026The Machines that Make the Machines
  3. Oct 6, 2026Introducing Mistral Large 4: Le chonk

Related