Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background.
Key points
- Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process.
- Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process.
- For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference.
- Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized image.Extensive experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation.
Sources (1)
- [1]Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:18 AM
Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background.
Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size
- Oct 4, 2026How corner is a corner case? Percentile control for highway scenario generation
- Sep 30, 2026ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images
- Sep 29, 2026Yzmblog/DMAD: DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
- Aug 10, 2026vllm-project/vllm v0.27.0