Backend-Agnostic Sparse Attention for Fast High-Resolution Visual Generation
To tackle these challenges, we propose BASA, a backend-agnostic sparse attention, which brings the best of both worlds: visual quality and practical acceleration.
Key points
- Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive.
- Window attention offers an efficient alternative, yet existing methods face a practical trade-off: partitioned window attention typically achieves computational efficiency consistent with its theoretical complexity.
- Fine-grained sliding-window attention effectively restores interactions across neighboring windows and improves visual quality.
- By introducing a structured window-shifting scheme across DiT blocks, we allow tokens divided by window boundaries in one layer to communicate in the following layers, thereby achieving global information exchange and eliminating window-induced visual artifacts.
Sources (1)
- [1]Backend-Agnostic Sparse Attention for Fast High-Resolution Visual GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:55 PM
To tackle these challenges, we propose BASA, a backend-agnostic sparse attention, which brings the best of both worlds: visual quality and practical acceleration.
Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive.
Extractive summary: sentences quoted from the sources.