RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing
We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms.
ProofPaper ↗
Key points
- Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene.
- Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues.
- We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references.
- These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.
Sources (1)
- [1]RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial RoutingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:16 AM
We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms.
Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 5, 2026LiquidAI/d1-omni-600M
- Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
- Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
- Sep 29, 2026How Diffusion Controller unifies and simplifies AI image generation
- Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation