Research

Papers and datasets worth knowing, ranked by significance and community attention.

Paper
Hugging Face Daily Papers2 sources3d ago

DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training

We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.

▲ 37 upvotesPaper
Paper
Apple Machine Learning Research2 sources5d ago

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation

In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.

▲ 71 upvotesPaper
Paper
Hugging Face Daily Papers2 sources5d ago

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.

▲ 156 upvotesPaper
Paper
Hugging Face Daily Papers2 sources4d ago

Q-Learning with Scalar Adjoint Matching

Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.

▲ 21 upvotesPaper
Paper
Hugging Face Daily Papers2 sources4d ago

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.

▲ 23 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization

We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.

▲ 8 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?

Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments.

▲ 135 upvotesPaper
Paper
Hugging Face Daily Papers2 sources4d ago

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

We introduce Δ-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills

We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy

In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

Agent Plasticity: Measuring Self-Improvement Through Experience

We introduce agent plasticity, the efficiency with which an agent converts experience into gains in future held-out performance.

Paper
Paper
Hugging Face Daily Papers3d ago

Opera: A Verbal Critic Framework for Long-horizon Coding Agents

We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

System Switch: When Should a Fast Decision Model Stop and Think?

Dual-process agents pair a fast policy with a slow deliberative model.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution

Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards

We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Safe Meta-Policy Design with Risk Control

We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents

We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition

Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata

In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

From Probabilities to Decisions: Search and Multi-Teacher Distillation with Jev

In bullet chess, a bot that places Jev's judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents

We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Cost-Aware Mixture-of-Experts Coordination for Model Markets

This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Executing Causal Structure Learning with Linear-Attention Transformers

We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation

We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity.

Paper
Paper
Hugging Face Daily Papers5d ago

A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning

We introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Shared and structured inputs undermine collective random choice by reasoning AI agents

Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Verification and Self-Improvement in Agentic AI: Foundations and Limits

Agentic AI systems can improve by searching longer, receiving additional support, or modifying how they propose and verify outputs.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

CausalDreamer: Learning Predictive World Models with Latent Disentanglement

We propose CausalDreamer, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents

We present a multi-aspect runtime-verification framework that decomposes a natural-language policy clause into a typed spatial/temporal/semantic triple over one canonical event stream, checks each aspect with its own monitoring specification, and fuses the verdicts through a four-valued algebra that carries provenance.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization

To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers

Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

AdaptEvo: Adaptive Agent Learning with Evolving Supervision

To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty

To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Measuring and Mitigating Solution Mode Collapse in RLVR

A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration

Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding

We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning

Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation

We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

MemoWM: How World Models Change What Agents Need to Remember

We formulate the problem of memory allocation conditioned on a world model and introduce MemoWM, a framework that uses shared predictions to compress retained information and reconstruct omitted content.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Policy Alignment: New Signals for Membership Auditing in On-Policy Distillation

In this paper, we propose Policy Alignment Membership Auditing (PAMA), a new auditing framework tailored for OPD.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Decoupling Exploration from Optimization in RLVR

Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Experience-Guided Initiation Search for Learned Skills in Skill Composition

We propose EVIS, an Experience-Guided and Behavior-Validated Initiation Search framework for discovering reliable initiation configurations under limited target interaction.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation

Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce agentic SElf-distilLation with environmental Feedback modeling (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs?

Paper