Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents.
Key points
- Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions.
- Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency.
- We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps.
- Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories.
Sources (1)
- [1]Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:48 PM
In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents.
Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions.
Extractive summary: sentences quoted from the sources.