Reward models
Also known as: reward model, reward modeling
15stories this week
15last 30 days
15all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceRubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-DistillationTo this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
- Oct 8, 2026 · Research paper · 2 sourcesDreamTrue: Action-Faithful Robot World Model with Counterfactual Post-TrainingWe present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
- Oct 8, 2026 · Research paper · 1 sourceMental-Models for Multi-Agent SystemsWe introduce mental-model-enabled agents, a framework that equips an agent with a latent mental model of its counterpart, allowing it to infer hidden beliefs, intentions, and likely reactions from the observed history and use these inferences to guide action selection.
- Oct 8, 2026 · Research paper · 1 sourceUXBench Pro: Benchmarking Personalized User Experience in Multi-Turn Dialogue InteractionsIn this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains.
- Oct 7, 2026 · Research paper · 1 sourceBudgeted Multi-Source Counterfactual Annotation for Off-Policy EvaluationOff-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation.
- Oct 7, 2026 · Research paper · 1 sourceI would rather quit NLP than read another paper like this: The rise of antithesis in NLP papersWe study the construction rather than in ACL papers from 2019, ACL-style arXiv papers from 2026, and papers written by GPT models from the same titles and abstracts.
- Oct 7, 2026 · Research paper · 1 sourceBoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token AggregationWe propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics.
- Oct 7, 2026 · Research paper · 1 sourceJudging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving CompressionWe introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction.
- Oct 7, 2026 · Research paper · 1 sourceWorld Potential Model: Pretrained World Knowledge as Progress PotentialsWe formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts.
- Oct 7, 2026 · Research paper · 1 sourceEfficient Best-of-N policy evaluation for inference-time alignmentBest-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model.
- Oct 6, 2026 · Research paper · 1 sourcePersonalize at Test Time: Learning User Preferences for Image GenerationWe propose an approach that learns personalized reward models directly from users' historical image preference pairs.
- Oct 6, 2026 · Research paper · 1 sourceVerify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving AgentsAs large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML).
- Oct 6, 2026 · Research paper · 1 sourceVariance-Optimal Off-Policy Evaluation with Conjunct Effect ModelingOff-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance.
- Oct 6, 2026 · Research paper · 1 sourceReinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with TransformersReinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses.
- Oct 6, 2026 · Research paper · 1 sourceSIGMA: Self-Improving Alignment Generalization from a Model SpecWe ask whether current models can improve their own safety alignment, and propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings.