PHBA: Prefix-State Hybrid Block Attention
In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context.
Key points
- Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval.
- Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window.
- We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors.
- Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.
Sources (1)
- [1]PHBA: Prefix-State Hybrid Block AttentionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:22 PM
In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context.
Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval.
Extractive summary: sentences quoted from the sources.