AION
Research paperLarge Language Models · Retrieval, RAG & Search · Efficiency & Inference1 source · Oct 6, 2026

PHBA: Prefix-State Hybrid Block Attention

In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context.

Key points

  • Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval.
  • Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window.
  • We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors.
  • Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.

Sources (1)

  • [1]PHBA: Prefix-State Hybrid Block Attention
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:22 PM
    In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context.
    Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval.

Extractive summary: sentences quoted from the sources.