Research
Papers and datasets worth knowing, ranked by significance and community attention.
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
PaperRISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.
Q-Learning with Scalar Adjoint Matching
Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?
Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments.
Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
We introduce Δ-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed.
ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories.
Agent Plasticity: Measuring Self-Improvement Through Experience
We introduce agent plasticity, the efficiency with which an agent converts experience into gains in future held-out performance.
Opera: A Verbal Critic Framework for Long-horizon Coding Agents
We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
System Switch: When Should a Fast Decision Model Stop and Think?
Dual-process agents pair a fast policy with a slow deliberative model.
Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution
Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
Safe Meta-Policy Design with Risk Control
We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression.
On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively.
Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition
Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging.
MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata
In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference.
From Probabilities to Decisions: Search and Multi-Teacher Distillation with Jev
In bullet chess, a bot that places Jev's judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots.
Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents
We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents.
Cost-Aware Mixture-of-Experts Coordination for Model Markets
This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.
Executing Causal Structure Learning with Linear-Attention Transformers
We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.
PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation
We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity.
A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning
We introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks.
Shared and structured inputs undermine collective random choice by reasoning AI agents
Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems.
Verification and Self-Improvement in Agentic AI: Foundations and Limits
Agentic AI systems can improve by searching longer, receiving additional support, or modifying how they propose and verify outputs.
CausalDreamer: Learning Predictive World Models with Latent Disentanglement
We propose CausalDreamer, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups.
SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents
We present a multi-aspect runtime-verification framework that decomposes a natural-language policy clause into a typed spatial/temporal/semantic triple over one canonical event stream, checks each aspect with its own monitoring specification, and fuses the verdicts through a four-valued algebra that carries provenance.
Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization
To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication.
Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers
Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses.
AdaptEvo: Adaptive Agent Learning with Evolving Supervision
To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics.
RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty
To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence.
Measuring and Mitigating Solution Mode Collapse in RLVR
A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces.
A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration
Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots.
CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding
We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning.
From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate.
World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task.
DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities.
MemoWM: How World Models Change What Agents Need to Remember
We formulate the problem of memory allocation conditioned on a world model and introduce MemoWM, a framework that uses shared predictions to compress retained information and reconstruct omitted content.
Policy Alignment: New Signals for Membership Auditing in On-Policy Distillation
In this paper, we propose Policy Alignment Membership Auditing (PAMA), a new auditing framework tailored for OPD.
Decoupling Exploration from Optimization in RLVR
Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints.
Experience-Guided Initiation Search for Learned Skills in Skill Composition
We propose EVIS, an Experience-Guided and Behavior-Validated Initiation Search framework for discovering reliable initiation configurations under limited target interaction.
Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation
Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce agentic SElf-distilLation with environmental Feedback modeling (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation.
When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better
In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs?