AION
Technique

RLVR

Also known as: reinforcement learning with verifiable rewards, verifiable rewards

10stories this week
12last 30 days
12all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
    We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
  2. Oct 8, 2026 · Research paper · 1 source
    RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty
    To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence.
  3. Oct 8, 2026 · Research paper · 1 source
    When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction
    We propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces.
  4. Oct 8, 2026 · Research paper · 1 source
    Measuring and Mitigating Solution Mode Collapse in RLVR
    A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces.
  5. Oct 7, 2026 · Research paper · 1 source
    SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
    Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
  6. Oct 7, 2026 · Research paper · 1 source
    VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
    Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),
  7. Oct 7, 2026 · Research paper · 1 source
    Decoupling Exploration from Optimization in RLVR
    Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints.
  8. Oct 7, 2026 · Research paper · 1 source
    Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
    In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents.
  9. Oct 7, 2026 · Research paper · 1 source
    RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation
    We introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction.
  10. Oct 6, 2026 · Research paper · 2 sources
    Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
    Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.
  11. Sep 28, 2026 · Research paper · 1 source
    Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
    Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze.
  12. Sep 28, 2026 · Research paper · 1 source
    ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
    To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an α-divergence variational objective and an exponential variance-control tilt.

Often appears with