PageWeaver: KV-Guided Query Unions for Sparse Attention
Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work.
Key points
- Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together.
- We present PageWeaver, an execution design that uses selected-page affinity to assemble query groups while preserving each query's original support and complete output ownership.
- A bounded GPU search produces query IDs, and an ID-aware two-CTA kernel consumes them without materializing reordered Q tensors or cross-page partial outputs.
- A direct KV-page union implementation provides a complementary design study of nonlocal reuse and reduction cost.
Sources (1)
- [1]PageWeaver: KV-Guided Query Unions for Sparse AttentionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:04 AM
Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work.
Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together.
Extractive summary: sentences quoted from the sources.