AION
Research paperImage, Video & 3D Generation1 source · Oct 6, 2026

Backend-Agnostic Sparse Attention for Fast High-Resolution Visual Generation

To tackle these challenges, we propose BASA, a backend-agnostic sparse attention, which brings the best of both worlds: visual quality and practical acceleration.

Key points

  • Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive.
  • Window attention offers an efficient alternative, yet existing methods face a practical trade-off: partitioned window attention typically achieves computational efficiency consistent with its theoretical complexity.
  • Fine-grained sliding-window attention effectively restores interactions across neighboring windows and improves visual quality.
  • By introducing a structured window-shifting scheme across DiT blocks, we allow tokens divided by window boundaries in one layer to communicate in the following layers, thereby achieving global information exchange and eliminating window-induced visual artifacts.

Sources (1)

  • [1]Backend-Agnostic Sparse Attention for Fast High-Resolution Visual Generation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:55 PM
    To tackle these challenges, we propose BASA, a backend-agnostic sparse attention, which brings the best of both worlds: visual quality and practical acceleration.
    Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive.

Extractive summary: sentences quoted from the sources.