Research

Papers and datasets worth knowing, ranked by significance and community attention.

Paper
Hugging Face Daily Papers2 sources3d ago

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.

▲ 52 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

WorldCast: Distributed Multiplayer World Models

We present WorldCast, a distributed multiplayer world model in which each player runs a local client comprising a video generator and a state model.

Paper
Paper
Apple Machine Learning Research3d ago

Normalizing Trajectory Models

We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training.

Paper
Hugging Face Daily Papers2 sources3d ago

LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.

▲ 24 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Reasoning-Informed Visual Editing

To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.

▲ 21 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs

OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint.

▲ 41 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.

▲ 8 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

VibeEdit: Image Editing with Canvas Instructions

We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.

▲ 18 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

AgentGarten: Code Worlds for Evolving Agents

We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments.

▲ 149 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching

To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Pumpire: Unified Benchmark for Metric Distance Estimation

We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

Q-Learning with Scalar Adjoint Matching

Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

From Prompting to Composing: A Spatial Canvas Interface for Poster Generation

We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

SpaceFlow: Locally Controllable 3D Generation

We present SpaceFlow, a training-free pipeline for locally controllable 3D generation from text descriptions and a collection of geometric primitives.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

The Lattice of Transition Laws

Diffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

CADFather Reconstructs Parametric CAD Programs Using Coordinated Tools

CADFather is an autonomous agentic system that coordinates complementary tools and a vision-language assistant to reconstruct parametric CAD models from 3D meshes without additional training.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent

To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's explicit requirements and its implied musical intent.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

SPW-Nav Streams Language-Guided Panoramic Video in Real Time

Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers

In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

From Video Clips to Creation Trajectory: Sora100K for AI-Native Video Creation

In this paper, we introduce Sora100K, a dataset that represents the AI-Native video creation workflow as a structured video creation trajectory.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation

To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Hybrid Cinematography: Previsualizing and Managing Hallucination Risk in Generative Video Reshooting

We present Hybrid Cinematography, a workflow that bridges physical capture and generative reshooting to manage hallucination risk while filmmakers can still act on it.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Equal Path Cost, Unequal Output Effects: Understanding Perturbation Propagation in Diffusion Models

To address this question, we develop a theoretical framework to investigate perturbation propagation, combining dynamical analysis of the sampling process with an information-theoretic characterization of output responses.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DataVista: Diagnosing Multimodal LLMs on Data Video Understanding

We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Learning to Retrieve: Internalizing Memory Retrieval for Video World Models

We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models

Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Parametric Trajectory Distillation for Few-Step Video Generation

We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

EchoDiST: Self-distillation-based joint learning for diffusion-conditioned echocardiographic myocardial motion estimation

We propose EchoDiST, a framework for unsupervised echocardiographic myocardial motion estimation that integrates self-distillation-based joint learning with a diffusion-conditioned motion estimation network.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

$C_4$-Equivariant Flow Matching on Anisotropic Power-Diagram Graphs for Microstructure Generation

We introduce a generative model for synthesising realistic polycrystalline microstructures using flow matching and graph neural networks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Finsler Flow Matching: Dynamics-Aware Geodesic Interpolation for Single-Snapshot Trajectory Inference

We introduce Finsler Flow Matching (FFM), a framework for learning continuous stochastic dynamics from discrete Markov transition graphs.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models

Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling

In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators

To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images

We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Diffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian Analysis

Despite their empirical success, why diffusion models overcome the bottlenecks of classical score-based samplers remains unclear.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

GRACE: Generation-aware latent compression for efficient video generation

To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

AuraLuxMuse: Adaptive Fusion Modeling for Aesthetic Stage Lighting Design with Music and Expert Guidance

We present AuraLuxMuse, a novel system for automated aesthetic stage lighting design that integrates expert knowledge, representation learning, and preference-adaptive modeling.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation

Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

ContiLNN: Mitigating Slice Sampling Discontinuity with Liquid Neural Networks for Medical Image Restoration

We introduce ContiLNN, which augments two-dimensional restoration backbones with bidirectional closed-form continuous-time (Bi-CfC) modules for cross-slice modeling while retaining in-plane feature extraction.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models

With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Connected Self Forcing: Beyond Local Learning in Video Autoregression

To stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Streaming-Aware Diffusion for Real-Time Video Super-Resolution via Cross-Step Attention

We propose a streaming-aware framework that adapts pretrained single-image latent diffusion models for efficient video super-resolution (VSR) by exploiting the sequential structure of video streams.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Towards Unified Evaluation of Prompt Enhancers for Video Generation

To address this gap, we introduce PEBench, the first unified benchmark for direct PE evaluation across text-to-video, image-to-video, and reference-to-video prompt enhancement.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Few-Step Generation via Data-Space Iteration

Flow matching has emerged as a scalable paradigm for training high-quality generative models, but sampling from the learned probability flow requires many network evaluations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution

We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

iCATS: Fast Video Generation via Interaction-Aware Sparse Attention and Timestep-Adaptive Sparsity

Training-free sparse attention offers a practical acceleration solution to Diffusion Transformers (DiTs) via reducing computations without fine-tuning.

Paper