Diffusion models
Also known as: DiT, diffusion model, diffusion transformer
64stories this week
66last 30 days
71all time
Timeline
- Oct 8, 2026 · Research paper · 2 sourcesLEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video GenerationGenerating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
- Oct 8, 2026 · Research paper · 1 sourceControl-Ready Uncertainty for Trajectory DiffusionWe introduce Score-Curvature for Online Precision Estimation (SCOPE), a lightweight module that augments diffusion trajectory models with control-ready uncertainty.
- Oct 8, 2026 · Research paper · 1 sourceAmbient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient LearningWe introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications.
- Oct 8, 2026 · Research paper · 1 sourceJust Weather Scoring: Efficient End-to-end Nowcasting with Distributional DiffusionWe introduce Just Weather Scoring (JWS), a single-stage, end-to-end diffusion model which addresses both issues by forecasting directly in radar space and enabling few-step generation.
- Oct 8, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.905-beta: Sandboxing is here!We're introducing Windows, Mac and Linux sandboxing in Unsloth!
- Oct 8, 2026 · Research paper · 1 sourceDiffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian AnalysisDespite their empirical success, why diffusion models overcome the bottlenecks of classical score-based samplers remains unclear.
- Oct 8, 2026 · Research paper · 1 sourceREACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA ModelsFlow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities.
- Oct 8, 2026 · Research paper · 1 sourceStreaming-Aware Diffusion for Real-Time Video Super-Resolution via Cross-Step AttentionWe propose a streaming-aware framework that adapts pretrained single-image latent diffusion models for efficient video super-resolution (VSR) by exploiting the sequential structure of video streams.
- Oct 8, 2026 · Research paper · 1 sourceHI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High ResolutionWe present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution.
- Oct 8, 2026 · Research paper · 1 sourceEarly Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic DenoisingWe show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization.
- Oct 8, 2026 · Research paper · 1 sourceStop My Dancing! Understanding, Detecting and Attributing Motion-Aware Deepfake VideosPose-guided diffusion models can now synthesize entire human figures in motion, spawning a new class of deepfakes: Motion Aware Deepfake (MAD) that have already reached hundreds of millions of viewers.
- Oct 8, 2026 · Research paper · 1 sourceConditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional TeacherCausal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation.
- Oct 8, 2026 · Research paper · 1 sourceWAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action ModelsWorld Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT).
- Oct 8, 2026 · Research paper · 1 sourceEqual Path Cost, Unequal Output Effects: Understanding Perturbation Propagation in Diffusion ModelsTo address this question, we develop a theoretical framework to investigate perturbation propagation, combining dynamical analysis of the sampling process with an information-theoretic characterization of output responses.
- Oct 8, 2026 · Research paper · 1 sourceCRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion DecodingLatent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces.
- Oct 8, 2026 · Research paper · 1 sourceBernoulli Flow Models: Self-Consistent Generative Modeling for Binary DataTo address this fundamental limitation and decouple the generative dynamics from fixed discrete time steps, we propose Bernoulli Flow Models (BFM).
- Oct 8, 2026 · Research paper · 1 sourceSample-Efficient Generative Conformal PredictionGenerative conformal prediction builds uncertainty sets from samples of a conditional generator, which are efficient only when the samples represent the response distribution well.
- Oct 8, 2026 · Research paper · 1 sourceMemorization and Malign Generalization in Conditional Diffusion Models with Random FeaturesConditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions.
- Oct 8, 2026 · Research paper · 1 sourceDynaTE: Accelerating Diffusion LLMs via Dynamic Token ExecutionDiffusion-based LLMs (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs by enabling bidirectional parallel refinement, alleviating the sequential decoding bottleneck of AR generation.
- Oct 8, 2026 · Research paper · 1 sourceAttributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion ModelsDiffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging.
- Oct 8, 2026 · Research paper · 2 sourcesThe Lattice of Transition LawsDiffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens.
- Oct 8, 2026 · Research paper · 1 sourceNo Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise ShapingAudio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.
- Oct 8, 2026 · Research paper · 1 sourceDiffusion Meta-Prompting and Steering for Generalizable Foundation Model AdaptationIn this paper, we introduce a Diffusion Meta-Prompt (DMP) model , a framework that models the distribution of learned prompts using diffusion models.
- Oct 8, 2026 · Research paper · 1 sourceTransforming Image Editors into Video EditorsIn this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation.
- Oct 7, 2026 · Research paper · 1 sourceEnabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion ModelsText-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference.
- Oct 7, 2026 · Research paper · 1 sourceGRACE: Generation-aware latent compression for efficient video generationTo address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT.
- Oct 7, 2026 · Research paper · 1 sourceKoopman Observers for Diffusion Acceleration: Correcting Feature Forecasts with Shallow MeasurementsWe introduce an observation-corrected Koopman framework for accelerating frozen diffusion models.
- Oct 7, 2026 · Research paper · 1 sourceReal-Time Joint Audio-Video Generation by Parallel Adapter CompositionDeploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap.
- Oct 7, 2026 · Research paper · 1 sourcePosition Forcing: Self-Conditioning 3D GenerationRecent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens.
- Oct 7, 2026 · Research paper · 1 sourceOrthoGen: A Generative Orthogonal Learner for Time-Varying TreatmentsEstimating conditional distributional potential outcomes (CDPOs) over time is important in medicine (e.g., to estimate patient-specific risks under different treatment sequences).