GRPODropout: Less is More for Online Reinforcement Learning Rollouts
Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement.
ProofPaper ↗
Key points
- We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates.
- To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages.
- To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection.
- This work provides insight into RL rollout usage: removing some rollouts can improve performance.
Sources (1)
- [1]GRPODropout: Less is More for Online Reinforcement Learning RolloutsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 12:38 PM
Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement.
We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
- Oct 8, 2026PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation
- Oct 8, 2026SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
- Oct 8, 2026Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
- Oct 7, 2026On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
- Oct 6, 2026huggingface/trl v1.14.2