Research
Papers and datasets worth knowing, ranked by significance and community attention.
Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
WorldCast: Distributed Multiplayer World Models
We present WorldCast, a distributed multiplayer world model in which each player runs a local client comprising a video generator and a state model.
PaperNormalizing Trajectory Models
We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training.
LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
Reasoning-Informed Visual Editing
To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.
OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint.
WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.
VibeEdit: Image Editing with Canvas Instructions
We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.
AgentGarten: Code Worlds for Evolving Agents
We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments.
Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes.
Pumpire: Unified Benchmark for Metric Distance Estimation
We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors.
Q-Learning with Scalar Adjoint Matching
Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.
SpaceFlow: Locally Controllable 3D Generation
We present SpaceFlow, a training-free pipeline for locally controllable 3D generation from text descriptions and a collection of geometric primitives.
The Lattice of Transition Laws
Diffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens.
Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.
CADFather Reconstructs Parametric CAD Programs Using Coordinated Tools
CADFather is an autonomous agentic system that coordinates complementary tools and a vision-language assistant to reconstruct parametric CAD models from 3D meshes without additional training.
MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent
To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's explicit requirements and its implied musical intent.
SPW-Nav Streams Language-Guided Panoramic Video in Real Time
Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.
Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.
From Video Clips to Creation Trajectory: Sora100K for AI-Native Video Creation
In this paper, we introduce Sora100K, a dataset that represents the AI-Native video creation workflow as a structured video creation trajectory.
Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation
To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
Hybrid Cinematography: Previsualizing and Managing Hallucination Risk in Generative Video Reshooting
We present Hybrid Cinematography, a workflow that bridges physical capture and generative reshooting to manage hallucination risk while filmmakers can still act on it.
Equal Path Cost, Unequal Output Effects: Understanding Perturbation Propagation in Diffusion Models
To address this question, we develop a theoretical framework to investigate perturbation propagation, combining dynamical analysis of the sampling process with an information-theoretic characterization of output responses.
DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains.
Learning to Retrieve: Internalizing Memory Retrieval for Video World Models
We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system.
Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models
Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging.
Parametric Trajectory Distillation for Few-Step Video Generation
We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path.
EchoDiST: Self-distillation-based joint learning for diffusion-conditioned echocardiographic myocardial motion estimation
We propose EchoDiST, a framework for unsupervised echocardiographic myocardial motion estimation that integrates self-distillation-based joint learning with a diffusion-conditioned motion estimation network.
$C_4$-Equivariant Flow Matching on Anisotropic Power-Diagram Graphs for Microstructure Generation
We introduce a generative model for synthesising realistic polycrystalline microstructures using flow matching and graph neural networks.
Finsler Flow Matching: Dynamics-Aware Geodesic Interpolation for Single-Snapshot Trajectory Inference
We introduce Finsler Flow Matching (FFM), a framework for learning continuous stochastic dynamics from discrete Markov transition graphs.
Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models
Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference.
No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.
VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling
In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data.
False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators
To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats.
GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images
We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images.
Diffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian Analysis
Despite their empirical success, why diffusion models overcome the bottlenecks of classical score-based samplers remains unclear.
GRACE: Generation-aware latent compression for efficient video generation
To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT.
AuraLuxMuse: Adaptive Fusion Modeling for Aesthetic Stage Lighting Design with Music and Expert Guidance
We present AuraLuxMuse, a novel system for automated aesthetic stage lighting design that integrates expert knowledge, representation learning, and preference-adaptive modeling.
WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation
Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency.
ContiLNN: Mitigating Slice Sampling Discontinuity with Liquid Neural Networks for Medical Image Restoration
We introduce ContiLNN, which augments two-dimensional restoration backbones with bidirectional closed-form continuous-time (Bi-CfC) modules for cross-slice modeling while retaining in-plane feature extraction.
Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models
With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).
Connected Self Forcing: Beyond Local Learning in Video Autoregression
To stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching.
Streaming-Aware Diffusion for Real-Time Video Super-Resolution via Cross-Step Attention
We propose a streaming-aware framework that adapts pretrained single-image latent diffusion models for efficient video super-resolution (VSR) by exploiting the sequential structure of video streams.
Towards Unified Evaluation of Prompt Enhancers for Video Generation
To address this gap, we introduce PEBench, the first unified benchmark for direct PE evaluation across text-to-video, image-to-video, and reference-to-video prompt enhancement.
Few-Step Generation via Data-Space Iteration
Flow matching has emerged as a scalable paradigm for training high-quality generative models, but sampling from the learned probability flow requires many network evaluations.
HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution
We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution.
iCATS: Fast Video Generation via Interaction-Aware Sparse Attention and Timestep-Adaptive Sparsity
Training-free sparse attention offers a practical acceleration solution to Diffusion Transformers (DiTs) via reducing computations without fine-tuning.