QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution
We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation.
Key points
- Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizing KV entries shared across queries.
- We observe substantial overlap in the KV entries selected by neighboring queries, creating opportunities for cross-query reuse.
- QUILT introduces Shift-and-Compare Set Decomposition (SCSD), which transforms irregular set operations into regular data-parallel primitives suitable for modern accelerators, and pipelines SCSD with attention computation to hide its overhead.
- We evaluate QUILT on LongBench using GLM-5.3 and DeepSeek-3.2 under both tensor and sequence parallelism.
Sources (1)
- [1]QUILT: Rethinking Sparse-Attention Prefill through Shared Query ExecutionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:58 AM
We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation.
Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizing KV entries shared across queries.
Extractive summary: sentences quoted from the sources.