ResearchResearch paperMultimodal Models · Large Language Models · Image, Video & 3D Generation1 source · Oct 7, 2026

ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count.

Key points

  • Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist.
  • We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two.
  • We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss.
  • Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace.

Sources (1)

  • [1]ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:03 AM
    Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count.
    Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  2. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  3. Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
  4. Aug 10, 2026vllm-project/vllm v0.27.0
  5. Jun 29, 2026vllm-project/vllm v0.24.0
  6. Jun 15, 2026vllm-project/vllm v0.23.0

Related