Research

Papers and datasets worth knowing, ranked by significance and community attention.

Paper
Hugging Face Daily Papers2 sources3d ago

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.

▲ 52 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

FearCaut-Qwen: Affective Steering in a Vision-Language Model Shifts the Decision Criterion for Hazard Assessment

Vision-language models (VLMs) show great potential for damage assessment after a disaster, but a recurring deficiency is that they are reluctant to declare a hazard; that is, recall is low even when overall accuracy appears adequate.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training

We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.

▲ 37 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement

We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy.

▲ 19 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

SuperNav: An Agentic Navigation System for Any Task in Any Scene

General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.

▲ 71 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

AgentGarten: Code Worlds for Evolving Agents

We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments.

▲ 149 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills

We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.

▲ 14 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces

Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces.

▲ 12 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching

To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes.

▲ 38 upvotesPaper
Paper
Hugging Face Daily Papers2 sources4d ago

Q-Learning with Scalar Adjoint Matching

Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

From Prompting to Composing: A Spatial Canvas Interface for Poster Generation

We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

BrickBench: Evaluating Agentic Brick Design

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models

Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation

To address these limitations, we formulate articulated asset reconstruction as programmatic modeling grounded in partial geometric evidence and introduce USDCraft, a framework in which a pretrained LLM writes and revises executable programs for simulation-ready articulated assets without task-specific training.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Predicting Cable Dynamics with Physical Attention Bias

Learned simulators for deformable linear objects (DLOs) such as cables have to predict the motion of cables they were not trained on and stay stable over long rollouts.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

CADFather Reconstructs Parametric CAD Programs Using Coordinated Tools

CADFather is an autonomous agentic system that coordinates complementary tools and a vision-language assistant to reconstruct parametric CAD models from 3D meshes without additional training.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy

In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC

We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

LiteNWM: Efficient Latent World Models for Onboard Visual Navigation in the Wild

We present LiteNWM, a latent navigation world model that shares visual encoding across candidates and jointly predicts their action-conditioned future representations at multiple horizons, while a learned scorer uses these predictions to select trajectories.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer

Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs).

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

SPW-Nav Streams Language-Guided Panoramic Video in Real Time

Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers

In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

CARE: Certifying Acceleration for Vision-Language-Action Inference

Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

On the estimation and validity of AI time horizons---a statistical look at the METR plot

On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation

To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

System Switch: When Should a Fast Decision Model Stop and Think?

Dual-process agents pair a fast policy with a slow deliberative model.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

CausalDreamer: Learning Predictive World Models with Latent Disentanglement

We propose CausalDreamer, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies

Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SpatialHarness: Test-Time Spatial Scaffolding for Fine Robotic Manipulation

We introduce SpatialHarness, a test-time embodied harness that provides test-time spatial scaffolding for fine robotic manipulation without policy fine-tuning or changes to the physical sensing setup.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?

A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models

Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

GLIO2: A GPU-Parallelized Tightly-Coupled LiDAR-Inertial-GNSS System for Robust and Real-Time Global Localization and Mapping

We propose GLIO2, a tightly-coupled LiDAR-Inertial-GNSS system whose GPU-parallel front-end jointly optimizes scan-to-multiscan LiDAR, IMU pre-integration, and raw GNSS measurements in a single sliding-window factor graph, sustaining real-time operation on edge hardware.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning

We study how a coding agent learns across a sequence of abstract reasoning tasks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Humanoid World Action Model With Joint State--Action Generation

We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

REACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA Models

Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models

World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding

We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations

Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Acting from Belief, Looking When Needed: A Bayesian Spatial World Model for Navigation under Intermittent Perception

We study navigation under intermittent perception: acting from an internal spatial belief and looking again only when execution needs a new observation, potentially freeing the shared sensor for other tasks between navigation observations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

RESETTLE: Robotic Recovery through Disagreement-Triggered Retrieval and Efficient Corrective Control

To address these challenges, we introduce RESETTLE(Robotic rEcovery through diSagrEement-Triggered reTrievaL and Efficient Corrective Control), a model-agnostic framework that provides computationally efficient recovery at the action-execution interface of frozen robot policies.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies

Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning

Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models

To address this, we propose WARP-VLA, a camera-view robust VLA for diverse wrist camera configurations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design

We present RFChipAgent, a first-of-its-kind multi-agent flow of large language model (LLM) agents for end-to-end analog/RF circuit design automation, in which AI agents collaboratively orchestrate the complete design flow under human supervision.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

LLM-IDEA: Identifiability-Driven Experimental Agent for Autonomous Discovery of Mechanistic World Models

We propose the Identifiability-Driven Experimental Agent (LLM-IDEA) for closed-loop discovery with an identifiability engine that returns a three-way plateau verdict: capability limit, resolvable within the design class, or certified exhausted.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Experience-Guided Initiation Search for Learned Skills in Skill Composition

We propose EVIS, an Experience-Guided and Behavior-Validated Initiation Search framework for discovering reliable initiation configurations under limited target interaction.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

SciExam for ENSO: Can AI Agents Build Climate Models?

The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations.

Paper