ResearchResearch paperRetrieval, RAG & Search1 source · Oct 6, 2026

Learning to Retrieve via Reinforcement Learning in Embedding Space

To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards.

Key points

  • Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance.
  • We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO).
  • As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation.
  • We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models.

Sources (1)

  • [1]Learning to Retrieve via Reinforcement Learning in Embedding Space
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:29 AM
    To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards.
    Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 5, 2026perplexity-ai/pplx-decider-v1.1-27b
  2. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  3. Oct 2, 2026alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
  4. Oct 1, 2026nvidia/PixelUMM
  5. Oct 1, 2026Barclays scales Claude to upgrade operations and improve client experience
  6. Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0

Related