Measuring and Mitigating Solution Mode Collapse in RLVR
A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces.
Key points
- Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered.
- We then use ModeBench to measure how solution diversity changes under RLVR post-training.
- We find that RLVR post-training concentrates probability onto fewer correct modes even as accuracy holds or improves, and moreover, that frontier models are already highly concentrated.
- We then introduce our solution, Re:Max, which stores one verified example per discovered mode in a replay buffer and trains on those stored modes uniformly.
Sources (1)
- [1]Measuring and Mitigating Solution Mode Collapse in RLVRarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:25 AM
A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces.
Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
- Oct 7, 2026VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
- Oct 7, 2026Decoupling Exploration from Optimization in RLVR
- Oct 7, 2026Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
- Oct 7, 2026RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation