Research

Papers and datasets worth knowing, ranked by significance and community attention.

Paper
Hugging Face Daily Papers2 sources3d ago

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.

▲ 52 upvotesPaper
Paper
Apple Machine Learning Research3d ago

Normalizing Trajectory Models

We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers

In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.

Paper
Paper
Google Research Blog12d ago

How Diffusion Controller unifies and simplifies AI image generation

We introduce Diffusion Controller, a lightweight "steering damper" network that precisely steers image generation to achieve significantly better prompt alignment.

Paper
Hugging Face Daily Papers2 sources3d ago

LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.

▲ 24 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Reasoning-Informed Visual Editing

To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.

▲ 21 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs

OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint.

▲ 41 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.

▲ 8 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

VibeEdit: Image Editing with Canvas Instructions

We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.

▲ 18 upvotesPaper
Paper
Hugging Face Daily Papers2 sources4d ago

Q-Learning with Scalar Adjoint Matching

Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

AgentGarten: Code Worlds for Evolving Agents

We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

CADFather Reconstructs Parametric CAD Programs Using Coordinated Tools

CADFather is an autonomous agentic system that coordinates complementary tools and a vision-language assistant to reconstruct parametric CAD models from 3D meshes without additional training.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching

To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

SPW-Nav Streams Language-Guided Panoramic Video in Real Time

Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Pumpire: Unified Benchmark for Metric Distance Estimation

We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

The Lattice of Transition Laws

Diffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens.

Paper
Paper
Hugging Face Daily Papers6d ago

MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers

Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

From Prompting to Composing: A Spatial Canvas Interface for Poster Generation

We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

SpaceFlow: Locally Controllable 3D Generation

We present SpaceFlow, a training-free pipeline for locally controllable 3D generation from text descriptions and a collection of geometric primitives.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent

To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's explicit requirements and its implied musical intent.

Paper
Paper
Hugging Face Daily Papers5d ago

SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models

Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

From Video Clips to Creation Trajectory: Sora100K for AI-Native Video Creation

In this paper, we introduce Sora100K, a dataset that represents the AI-Native video creation workflow as a structured video creation trajectory.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Equal Path Cost, Unequal Output Effects: Understanding Perturbation Propagation in Diffusion Models

To address this question, we develop a theoretical framework to investigate perturbation propagation, combining dynamical analysis of the sampling process with an information-theoretic characterization of output responses.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models

Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

WorldCast: Distributed Multiplayer World Models

We present WorldCast, a distributed multiplayer world model in which each player runs a local client comprising a video generator and a state model.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers

In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models

Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Learning to Retrieve: Internalizing Memory Retrieval for Video World Models

We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

GRACE: Generation-aware latent compression for efficient video generation

To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation

To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Hybrid Cinematography: Previsualizing and Managing Hallucination Risk in Generative Video Reshooting

We present Hybrid Cinematography, a workflow that bridges physical capture and generative reshooting to manage hallucination risk while filmmakers can still act on it.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Think Before You Paint: Recursive Latent Reasoning for Diffusion Models

We propose Painter-Thinker (PaTh): a small recursive network (the Thinker) reasons over a grid of learned tokens that encode the noisy image and the conditioning, refines a latent state within every denoising step, and steers a frozen diffusion model (the Painter) through ControlNet adapters.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Parametric Trajectory Distillation for Few-Step Video Generation

We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

EchoDiST: Self-distillation-based joint learning for diffusion-conditioned echocardiographic myocardial motion estimation

We propose EchoDiST, a framework for unsupervised echocardiographic myocardial motion estimation that integrates self-distillation-based joint learning with a diffusion-conditioned motion estimation network.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DataVista: Diagnosing Multimodal LLMs on Data Video Understanding

We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map

We present UltraWorld, a self-distillation recipe that transfers priors from clinical ultrasound videos into interactive world models without real action annotations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models

With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

SNR-Gated LSTM-Conditioned Diffusion Model for MIMO Channel Estimation

Accurate and low latency channel estimation is critical for modern MIMO systems, particularly under mobility, where channels exhibit structured sparsity and strong temporal correlation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Steering Diffusion Models to Rare Events with Sequential Monte Carlo

In this paper, we present Diffusion Importance Sampling of Rare Events or DireSMC, a sequential Monte Carlo scheme that guides a population of weighted samples towards the rare event, giving access not only to samples but also to a calibrated estimate of its probability.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators

To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

$C_4$-Equivariant Flow Matching on Anisotropic Power-Diagram Graphs for Microstructure Generation

We introduce a generative model for synthesising realistic polycrystalline microstructures using flow matching and graph neural networks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Controllable Crowd Generation through World-Model Planning

To address this limitation, we propose Ctrl-CWM, a multi-agent Controllable Crowd World Model that integrates crowd generation and run-time control.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Finsler Flow Matching: Dynamics-Aware Geodesic Interpolation for Single-Snapshot Trajectory Inference

We introduce Finsler Flow Matching (FFM), a framework for learning continuous stochastic dynamics from discrete Markov transition graphs.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images

We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation

We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation.

Paper