Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval
We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval.
Key points
- Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity.
- ARR alternates between retrieving an item and updating the query embedding, then uses the final embedding to rank the collection.
- ARR demonstrates strong retrieval performance on both in-domain and zero-shot benchmarks, outperforming the compared baselines on average.
- Further analyses show that feedback improves retrieval at inference time and that training with feedback also improves the initial query embedding, before any item is observed.
Sources (1)
- [1]Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal RetrievalarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 10:39 AM
We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval.
Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
- Oct 8, 2026SuperNav: An Agentic Navigation System for Any Task in Any Scene
- Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
- Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching