A Systematic Study of Semantic ID Spaces for Generative Information Retrieval
Generative Information Retrieval (GIR) has emerged as a transformative paradigm, shifting document retrieval from a traditional "retrieve-and-rank" workflow to sequence-to-sequence generation, where a model directly predicts document identifiers (DocIDs).
ProofPaper ↗
Key points
- While the semantic design of these DocIDs is known to be critical for performance, a fundamental question remains under-explored: what makes a good DocID?
- In this work, we address this challenge by presenting a comprehensive study on the properties, metrics, and trade-offs that define effective numerical DocIDs.
- Specifically, our contributions are threefold: First, we propose a unified framework that unifies Product Quantization (PQ) and Residual Quantization (RQ), and their hybrid variants within a single design space.
- Through extensive experiments on MS MARCO 300K and NQ320K, we analyze how these structural properties influence retrieval effectiveness.
Sources (1)
- [1]A Systematic Study of Semantic ID Spaces for Generative Information RetrievalarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:33 PM
Generative Information Retrieval (GIR) has emerged as a transformative paradigm, shifting document retrieval from a traditional "retrieve-and-rank" workflow to sequence-to-sequence generation, where a model directly predicts document identifiers (DocIDs).
While the semantic design of these DocIDs is known to be critical for performance, a fundamental question remains under-explored: what makes a good DocID?
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
- Sep 22, 2026vllm-project/vllm v0.30.0
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jul 11, 2026vllm-project/vllm v0.25.0
- Jun 29, 2026vllm-project/vllm v0.24.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation