RLVR
Also known as: reinforcement learning with verifiable rewards, verifiable rewards
10stories this week
12last 30 days
12all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceResidual Advantage: Student-Relative Teacher Guidance for RL with Verifiable RewardsWe propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
- Oct 8, 2026 · Research paper · 1 sourceRL-ARC: Calibrating Large Reasoning Models via Reasoning-guided UncertaintyTo this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence.
- Oct 8, 2026 · Research paper · 1 sourceWhen Interfaces Speak: Data-Aware Generative UI Harness for Active InteractionWe propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces.
- Oct 8, 2026 · Research paper · 1 sourceMeasuring and Mitigating Solution Mode Collapse in RLVRA language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces.
- Oct 7, 2026 · Research paper · 1 sourceSPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level AdvantagesVision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
- Oct 7, 2026 · Research paper · 1 sourceVICO: Visual Environments Co-Evolving for Vision-Language Model ReasoningReinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),
- Oct 7, 2026 · Research paper · 1 sourceDecoupling Exploration from Optimization in RLVRModern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints.
- Oct 7, 2026 · Research paper · 1 sourceBeyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search AgentsIn this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents.
- Oct 7, 2026 · Research paper · 1 sourceRewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward AdaptationWe introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction.
- Oct 6, 2026 · Research paper · 2 sourcesSelf-Retrospection Distillation: Turning Post-hoc Experiences into Prior ForesightReinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.
- Sep 28, 2026 · Research paper · 1 sourceLearning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable VectorsReinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze.
- Sep 28, 2026 · Research paper · 1 sourceReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy LearningTo address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an α-divergence variational objective and an exponential variance-control tilt.