Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers
Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear.
ProofPaper ↗
Key points
- In this paper, we systematically examine retrievers across architectures and find that their performance is highly sensitive to modality composition.
- As image documents are progressively replaced with semantically corresponding text representations, retrieval performance follows a pronounced V-shaped curve, remaining strong on single-modality corpora but degrading substantially when modalities coexist.
- To mitigate this bias, we introduce Trident, which constructs text, image, and fused text-image views of each document as co-equal positives and jointly optimizes relevance discrimination and positive-view balance through Multi-Positive View InfoNCE.
- Experiments across visual document and natural image benchmarks show that trident improves mixed-modality retrieval on both CLIP-based and VLM-based architectures, reduces sensitivity to modality composition and text distractors, and increases average single-modality retrieval performance.
Sources (1)
- [1]Chaos in the Text: Revealing the Modality Preference in Mixed-Modality RetrieversHugging Face Daily Papers · Oct 8, 12:00 AM
Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear.
In this paper, we systematically examine retrievers across architectures and find that their performance is highly sensitive to modality composition.
Extractive summary: sentences quoted from the sources.