AION
Research paperReinforcement Learning1 source · Oct 8, 2026

Measuring and Mitigating Solution Mode Collapse in RLVR

A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces.

Key points

  • Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered.
  • We then use ModeBench to measure how solution diversity changes under RLVR post-training.
  • We find that RLVR post-training concentrates probability onto fewer correct modes even as accuracy holds or improves, and moreover, that frontier models are already highly concentrated.
  • We then introduce our solution, Re:Max, which stores one verified example per discovered mode in a replay buffer and trains on those stored modes uniformly.

Sources (1)

  • [1]Measuring and Mitigating Solution Mode Collapse in RLVR
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:25 AM
    A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces.
    Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
  2. Oct 7, 2026VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
  3. Oct 7, 2026Decoupling Exploration from Optimization in RLVR
  4. Oct 7, 2026Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
  5. Oct 7, 2026RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation

Related