Technique

Reward models

Also known as: reward model, reward modeling

15stories this week
15last 30 days
15all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation
    To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
  2. Oct 8, 2026 · Research paper · 2 sources
    DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
    We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
  3. Oct 8, 2026 · Research paper · 1 source
    Mental-Models for Multi-Agent Systems
    We introduce mental-model-enabled agents, a framework that equips an agent with a latent mental model of its counterpart, allowing it to infer hidden beliefs, intentions, and likely reactions from the observed history and use these inferences to guide action selection.
  4. Oct 8, 2026 · Research paper · 1 source
    UXBench Pro: Benchmarking Personalized User Experience in Multi-Turn Dialogue Interactions
    In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains.
  5. Oct 7, 2026 · Research paper · 1 source
    Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation
    Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation.
  6. Oct 7, 2026 · Research paper · 1 source
    I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers
    We study the construction rather than in ACL papers from 2019, ACL-style arXiv papers from 2026, and papers written by GPT models from the same titles and abstracts.
  7. Oct 7, 2026 · Research paper · 1 source
    BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
    We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics.
  8. Oct 7, 2026 · Research paper · 1 source
    Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression
    We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction.
  9. Oct 7, 2026 · Research paper · 1 source
    World Potential Model: Pretrained World Knowledge as Progress Potentials
    We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts.
  10. Oct 7, 2026 · Research paper · 1 source
    Efficient Best-of-N policy evaluation for inference-time alignment
    Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model.
  11. Oct 6, 2026 · Research paper · 1 source
    Personalize at Test Time: Learning User Preferences for Image Generation
    We propose an approach that learns personalized reward models directly from users' historical image preference pairs.
  12. Oct 6, 2026 · Research paper · 1 source
    Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
    As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML).
  13. Oct 6, 2026 · Research paper · 1 source
    Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling
    Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance.
  14. Oct 6, 2026 · Research paper · 1 source
    Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers
    Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses.
  15. Oct 6, 2026 · Research paper · 1 source
    SIGMA: Self-Improving Alignment Generalization from a Model Spec
    We ask whether current models can improve their own safety alignment, and propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings.

Often appears with