ResearchResearch paperMultimodal Models · Retrieval, RAG & Search · Interpretability1 source · Oct 8, 2026

Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers

Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear.

Key points

  • In this paper, we systematically examine retrievers across architectures and find that their performance is highly sensitive to modality composition.
  • As image documents are progressively replaced with semantically corresponding text representations, retrieval performance follows a pronounced V-shaped curve, remaining strong on single-modality corpora but degrading substantially when modalities coexist.
  • To mitigate this bias, we introduce Trident, which constructs text, image, and fused text-image views of each document as co-equal positives and jointly optimizes relevance discrimination and positive-view balance through Multi-Positive View InfoNCE.
  • Experiments across visual document and natural image benchmarks show that trident improves mixed-modality retrieval on both CLIP-based and VLM-based architectures, reduces sensitivity to modality composition and text distractors, and increases average single-modality retrieval performance.

Sources (1)

  • [1]Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers
    Hugging Face Daily Papers · Oct 8, 12:00 AM
    Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear.
    In this paper, we systematically examine retrievers across architectures and find that their performance is highly sensitive to modality composition.

Extractive summary: sentences quoted from the sources.

Related