Research
Papers and datasets worth knowing, ranked by significance and community attention.
Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
FearCaut-Qwen: Affective Steering in a Vision-Language Model Shifts the Decision Criterion for Hazard Assessment
Vision-language models (VLMs) show great potential for damage assessment after a disaster, but a recurring deficiency is that they are reluctant to declare a hazard; that is, recall is low even when overall accuracy appears adequate.
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
SuperNav: An Agentic Navigation System for Any Task in Any Scene
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.
Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy.
Q-Learning with Scalar Adjoint Matching
Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
AgentGarten: Code Worlds for Evolving Agents
We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments.
SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces.
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.
CADFather Reconstructs Parametric CAD Programs Using Coordinated Tools
CADFather is an autonomous agentic system that coordinates complementary tools and a vision-language assistant to reconstruct parametric CAD models from 3D meshes without additional training.
ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models
Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.
Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes.
SPW-Nav Streams Language-Guided Panoramic Video in Real Time
Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.
CARE: Certifying Acceleration for Vision-Language-Action Inference
Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success.
USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation
To address these limitations, we formulate articulated asset reconstruction as programmatic modeling grounded in partial geometric evidence and introduce USDCraft, a framework in which a pretrained LLM writes and revises executable programs for simulation-ready articulated assets without task-specific training.
UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories.
From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.
BrickBench: Evaluating Agentic Brick Design
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design.
System Switch: When Should a Fast Decision Model Stop and Think?
Dual-process agents pair a fast policy with a slow deliberative model.
Predicting Cable Dynamics with Physical Attention Bias
Learned simulators for deformable linear objects (DLOs) such as cables have to predict the motion of cables they were not trained on and stay stable over long rollouts.
A Stevens's Power Law Check-up of GPT-5.5's Implicit Reading of Visual Encoding
We adapt Stevens's power law to measure the implicit ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models.
Spatial Latent Reasoning for Embodied Reference Understanding
We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states.
MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata
In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference.
Evaluating the Transfer of Co-Evolved Communication from 2D to 3D Simulation
This work examines the transfer of a co-evolved communication mechanism between two robotic agents from a discrete two-dimensional (2D) simulator to a three-dimensional simulator with real physics (3D).
Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents
We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents.
Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs).
Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.
Design of a Fully Actuated 4-DOF Robotic Finger With Joint-Specific Hybrid Remote Actuation
This paper presents a fully actuated 4-DOF robotic finger using a joint-specific hybrid remote-actuation architecture.
RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
We present RFChipAgent, a first-of-its-kind multi-agent flow of large language model (LLM) agents for end-to-end analog/RF circuit design automation, in which AI agents collaboratively orchestrate the complete design flow under human supervision.
SciExam for ENSO: Can AI Agents Build Climate Models?
The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations.
How Do Transformers Learn to Represent Symmetries?
Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning.
Fully Interpretable Minimal Transformers: From Geometry to Algorithm
We present a framework for building and interpreting minimal transformer models.
Shared and structured inputs undermine collective random choice by reasoning AI agents
Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems.
Verification and Self-Improvement in Agentic AI: Foundations and Limits
Agentic AI systems can improve by searching longer, receiving additional support, or modifying how they propose and verify outputs.
Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning
We study how a coding agent learns across a sequence of abstract reasoning tasks.
WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models
World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT).
SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning
Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation.
LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC
We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes.
LiteNWM: Efficient Latent World Models for Onboard Visual Navigation in the Wild
We present LiteNWM, a latent navigation world model that shares visual encoding across candidates and jointly predicts their action-conditioned future representations at multiple horizons, while a learned scorer uses these predictions to select trajectories.
CausalDreamer: Learning Predictive World Models with Latent Disentanglement
We propose CausalDreamer, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups.
Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
We introduce spatial induction heads, two-layer gather-and-match circuits in which the first layer reconstructs the relevant spatial neighborhood and the second matches the resulting configuration against earlier occurrences.
Acting from Belief, Looking When Needed: A Bayesian Spatial World Model for Navigation under Intermittent Perception
We study navigation under intermittent perception: acting from an internal spatial belief and looking again only when execution needs a new observation, potentially freeing the shared sensor for other tasks between navigation observations.
Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving.
LLM-IDEA: Identifiability-Driven Experimental Agent for Autonomous Discovery of Mechanistic World Models
We propose the Identifiability-Driven Experimental Agent (LLM-IDEA) for closed-loop discovery with an identifiability engine that returns a three-way plateau verdict: capability limit, resolvable within the design class, or certified exhausted.
On the estimation and validity of AI time horizons---a statistical look at the METR plot
On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time.
VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation
To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning.
PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies
Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation.