Disentangling Paradigm, Identifier, and Decoding in Generative Retrieval
Generative retrieval trains a language model to generate the identifier of a relevant document.
ProofPaper ↗
Key points
- Recent work replaces the autoregressive decoder with diffusion, but changes identifiers, training recipe and decoding at once, so differences cannot be credited to the paradigm.
- On NQ320K and MS300K, we train autoregressive, masked-diffusion and block-diffusion models with residual-quantised, product-quantised and random identifiers.
- Our reference diffusion decoding, generate-and-match, generates an identifier, then retrieves the closest corpus identifiers.
- We test one-pass scoring to decode diffusion retrievers: the model reads a fully masked identifier once, and each document is scored by its codes' probabilities.
Sources (1)
- [1]Disentangling Paradigm, Identifier, and Decoding in Generative RetrievalarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:23 PM
Generative retrieval trains a language model to generate the identifier of a relevant document.
Recent work replaces the autoregressive decoder with diffusion, but changes identifiers, training recipe and decoding at once, so differences cannot be credited to the paradigm.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Steering Diffusion Models to Rare Events with Sequential Monte Carlo
- Oct 6, 2026Enhancing Diffusion Language Models with Autoregressive Post-Training Weights
- Oct 6, 2026Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
- Oct 6, 2026Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size
- Oct 4, 2026How corner is a corner case? Percentile control for highway scenario generation
- Aug 10, 2026vllm-project/vllm v0.27.0