AION
Technique

GRPO

Also known as: group relative policy optimization

32stories this week
32last 30 days
32all time

Timeline

  1. Oct 8, 2026 · Open-source release · 1 source
    huggingface/trl v1.15.0
    SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
  2. Oct 8, 2026 · Research paper · 2 sources
    Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
    Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.
  3. Oct 8, 2026 · Research paper · 1 source
    ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation
    We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks.
  4. Oct 8, 2026 · Research paper · 1 source
    GRPODropout: Less is More for Online Reinforcement Learning Rollouts
    Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement.
  5. Oct 8, 2026 · Research paper · 1 source
    ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry
    We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.
  6. Oct 8, 2026 · Research paper · 1 source
    Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
    We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
  7. Oct 8, 2026 · Research paper · 1 source
    Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization
    To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication.
  8. Oct 8, 2026 · Research paper · 1 source
    Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
    Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
  9. Oct 8, 2026 · Research paper · 1 source
    Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation
    Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce agentic SElf-distilLation with environmental Feedback modeling (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation.
  10. Oct 8, 2026 · Research paper · 1 source
    From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning
    To address these limitations, we propose Evolving Functional REpertoires (EFRE), which replaces a single prompt with a repertoire of functions that evolves as new tasks arrive: compatible updates refine existing functions, while conflicting updates trigger the emergence of new ones.
  11. Oct 8, 2026 · Research paper · 2 sources
    SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
    Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces.
  12. Oct 8, 2026 · Research paper · 1 source
    AdaptEvo: Adaptive Agent Learning with Evolving Supervision
    To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics.
  13. Oct 8, 2026 · Research paper · 1 source
    PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation
    We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity.
  14. Oct 8, 2026 · Research paper · 1 source
    Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL
    A verified reference solution provides a correct trajectory for training a reasoning model.
  15. Oct 7, 2026 · Research paper · 1 source
    RH-Detect: A Unified Benchmark for Reward Hacking Detection
    We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema.
  16. Oct 7, 2026 · Research paper · 1 source
    StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents
    We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty.
  17. Oct 7, 2026 · Research paper · 1 source
    SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
    Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
  18. Oct 7, 2026 · Research paper · 1 source
    Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
    We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
  19. Oct 7, 2026 · Research paper · 1 source
    On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
    We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively.
  20. Oct 7, 2026 · Research paper · 1 source
    RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
    In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved.
  21. Oct 7, 2026 · Research paper · 1 source
    Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL
    When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it?
  22. Oct 7, 2026 · Research paper · 1 source
    Learning to Accumulate Knowledge with Mutual Information
    Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories.
  23. Oct 7, 2026 · Research paper · 1 source
    MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection
    In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues.
  24. Oct 7, 2026 · Research paper · 1 source
    BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
    We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics.
  25. Oct 7, 2026 · Research paper · 1 source
    World Potential Model: Pretrained World Knowledge as Progress Potentials
    We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts.
  26. Oct 7, 2026 · Research paper · 1 source
    Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
    We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals.
  27. Oct 6, 2026 · Open-source release · 1 source
    huggingface/trl v1.14.2
    Patch release fixing two cases of silently wrong training and three crashes.
  28. Oct 6, 2026 · Research paper · 1 source
    Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
    As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML).
  29. Oct 6, 2026 · Research paper · 1 source
    On KL-Regularized Policy Optimization
    We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler.
  30. Oct 6, 2026 · Research paper · 1 source
    AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails
    We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adaptive reasoning-budget allocation.

Often appears with