GRPO
Also known as: group relative policy optimization
32stories this week
32last 30 days
32all time
Timeline
- Oct 8, 2026 · Open-source release · 1 sourcehuggingface/trl v1.15.0SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
- Oct 8, 2026 · Research paper · 2 sourcesDistilling Routed 3D Privilege for Spatial Reasoning in Vision-Language ModelsSpatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.
- Oct 8, 2026 · Research paper · 1 sourceContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression SegmentationWe introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks.
- Oct 8, 2026 · Research paper · 1 sourceGRPODropout: Less is More for Online Reinforcement Learning RolloutsReinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement.
- Oct 8, 2026 · Research paper · 1 sourceReTeach: Building a Self-Teacher through Multi-Round Reflection and RetryWe introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.
- Oct 8, 2026 · Research paper · 1 sourceResidual Advantage: Student-Relative Teacher Guidance for RL with Verifiable RewardsWe propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
- Oct 8, 2026 · Research paper · 1 sourceFed-GRPO: Reward-Signal-Driven Federated Group Relative Policy OptimizationTo this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication.
- Oct 8, 2026 · Research paper · 1 sourceBeyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-SpeechNatural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
- Oct 8, 2026 · Research paper · 1 sourceEnvironmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-DistillationGiven that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce agentic SElf-distilLation with environmental Feedback modeling (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation.
- Oct 8, 2026 · Research paper · 1 sourceFrom a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual LearningTo address these limitations, we propose Evolving Functional REpertoires (EFRE), which replaces a single prompt with a repertoire of functions that evolves as new tasks arrive: compatible updates refine existing functions, while conflicting updates trigger the emergence of new ones.
- Oct 8, 2026 · Research paper · 2 sourcesSpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent TracesSpatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces.
- Oct 8, 2026 · Research paper · 1 sourceAdaptEvo: Adaptive Agent Learning with Evolving SupervisionTo address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics.
- Oct 8, 2026 · Research paper · 1 sourcePIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot DistillationWe propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity.
- Oct 8, 2026 · Research paper · 1 sourceBalancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RLA verified reference solution provides a correct trajectory for training a reasoning model.
- Oct 7, 2026 · Research paper · 1 sourceRH-Detect: A Unified Benchmark for Reward Hacking DetectionWe present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema.
- Oct 7, 2026 · Research paper · 1 sourceStoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator AgentsWe introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty.
- Oct 7, 2026 · Research paper · 1 sourceSPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level AdvantagesVision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
- Oct 7, 2026 · Research paper · 1 sourceConversational Voice Aesthetic Model with Reinforcement Learning from Human ListenersWe introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
- Oct 7, 2026 · Research paper · 1 sourceOn the Clock: Towards Punctual and Productive Time-Budgeted AI AgentsWe study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively.
- Oct 7, 2026 · Research paper · 1 sourceRECAST: Learning to Compute the Right Context through Adaptive Evidence RoutingIn this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved.
- Oct 7, 2026 · Research paper · 1 sourceWhich Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RLWhen reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it?
- Oct 7, 2026 · Research paper · 1 sourceLearning to Accumulate Knowledge with Mutual InformationTherefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories.
- Oct 7, 2026 · Research paper · 1 sourceMSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake DetectionIn this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues.
- Oct 7, 2026 · Research paper · 1 sourceBoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token AggregationWe propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics.
- Oct 7, 2026 · Research paper · 1 sourceWorld Potential Model: Pretrained World Knowledge as Progress PotentialsWe formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts.
- Oct 7, 2026 · Research paper · 1 sourceVisual Jev Rewards: Reference-Bound Verification for Multi-Subject Image GenerationWe present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals.
- Oct 6, 2026 · Open-source release · 1 sourcehuggingface/trl v1.14.2Patch release fixing two cases of silently wrong training and three crashes.
- Oct 6, 2026 · Research paper · 1 sourceVerify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving AgentsAs large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML).
- Oct 6, 2026 · Research paper · 1 sourceOn KL-Regularized Policy OptimizationWe propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler.
- Oct 6, 2026 · Research paper · 1 sourceAdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge GuardrailsWe present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adaptive reasoning-budget allocation.