UNREAL: Unifying Retrieval and Long-Context with a Single Model
We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference.
Key points
- Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus.
- UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations.
- On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems.
- Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.
Sources (1)
- [1]UNREAL: Unifying Retrieval and Long-Context with a Single ModelarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 02:48 PM
We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference.
Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus.
Extractive summary: sentences quoted from the sources.