Research
Papers and datasets worth knowing, ranked by significance and community attention.
Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
PaperNormalizing Trajectory Models
We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training.
Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.
How Diffusion Controller unifies and simplifies AI image generation
We introduce Diffusion Controller, a lightweight "steering damper" network that precisely steers image generation to achieve significantly better prompt alignment.
LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
Reasoning-Informed Visual Editing
To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.
OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint.
WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.
VibeEdit: Image Editing with Canvas Instructions
We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.
Q-Learning with Scalar Adjoint Matching
Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
AgentGarten: Code Worlds for Evolving Agents
We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments.
Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.
CADFather Reconstructs Parametric CAD Programs Using Coordinated Tools
CADFather is an autonomous agentic system that coordinates complementary tools and a vision-language assistant to reconstruct parametric CAD models from 3D meshes without additional training.
Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes.
SPW-Nav Streams Language-Guided Panoramic Video in Real Time
Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.
Pumpire: Unified Benchmark for Metric Distance Estimation
We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors.
The Lattice of Transition Laws
Diffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens.
MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation.
From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.
SpaceFlow: Locally Controllable 3D Generation
We present SpaceFlow, a training-free pipeline for locally controllable 3D generation from text descriptions and a collection of geometric primitives.
MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent
To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's explicit requirements and its implied musical intent.
SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass.
From Video Clips to Creation Trajectory: Sora100K for AI-Native Video Creation
In this paper, we introduce Sora100K, a dataset that represents the AI-Native video creation workflow as a structured video creation trajectory.
Equal Path Cost, Unequal Output Effects: Understanding Perturbation Propagation in Diffusion Models
To address this question, we develop a theoretical framework to investigate perturbation propagation, combining dynamical analysis of the sampling process with an information-theoretic characterization of output responses.
Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models
Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging.
WorldCast: Distributed Multiplayer World Models
We present WorldCast, a distributed multiplayer world model in which each player runs a local client comprising a video generator and a state model.
Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers
In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT.
Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models
Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference.
Learning to Retrieve: Internalizing Memory Retrieval for Video World Models
We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system.
GRACE: Generation-aware latent compression for efficient video generation
To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT.
Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation
To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
Hybrid Cinematography: Previsualizing and Managing Hallucination Risk in Generative Video Reshooting
We present Hybrid Cinematography, a workflow that bridges physical capture and generative reshooting to manage hallucination risk while filmmakers can still act on it.
Think Before You Paint: Recursive Latent Reasoning for Diffusion Models
We propose Painter-Thinker (PaTh): a small recursive network (the Thinker) reasons over a grid of learned tokens that encode the noisy image and the conditioning, refines a latent state within every denoising step, and steers a frozen diffusion model (the Painter) through ControlNet adapters.
ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count.
Parametric Trajectory Distillation for Few-Step Video Generation
We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path.
EchoDiST: Self-distillation-based joint learning for diffusion-conditioned echocardiographic myocardial motion estimation
We propose EchoDiST, a framework for unsupervised echocardiographic myocardial motion estimation that integrates self-distillation-based joint learning with a diffusion-conditioned motion estimation network.
DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains.
No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.
UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map
We present UltraWorld, a self-distillation recipe that transfers priors from clinical ultrasound videos into interactive world models without real action annotations.
Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models
With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).
SNR-Gated LSTM-Conditioned Diffusion Model for MIMO Channel Estimation
Accurate and low latency channel estimation is critical for modern MIMO systems, particularly under mobility, where channels exhibit structured sparsity and strong temporal correlation.
Steering Diffusion Models to Rare Events with Sequential Monte Carlo
In this paper, we present Diffusion Importance Sampling of Rare Events or DireSMC, a sequential Monte Carlo scheme that guides a population of weighted samples towards the rare event, giving access not only to samples but also to a calibrated estimate of its probability.
False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators
To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats.
$C_4$-Equivariant Flow Matching on Anisotropic Power-Diagram Graphs for Microstructure Generation
We introduce a generative model for synthesising realistic polycrystalline microstructures using flow matching and graph neural networks.
Controllable Crowd Generation through World-Model Planning
To address this limitation, we propose Ctrl-CWM, a multi-agent Controllable Crowd World Model that integrates crowd generation and run-time control.
Finsler Flow Matching: Dynamics-Aware Geodesic Interpolation for Single-Snapshot Trajectory Inference
We introduce Finsler Flow Matching (FFM), a framework for learning continuous stochastic dynamics from discrete Markov transition graphs.
GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images
We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images.
QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation
We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation.