Local Content-Style Control for Diffusion-based Image Stylization
Image stylization with latent-diffusion models entangles two independently refined axes: what a region depicts and how it is depicted.
ProofPaper ↗
Key points
- Such pipelines expose only global controls, yet professional retouching demands deliberate, region-specific control.
- We lift two conditioning weights already present in a ControlNet + IP-Adapter stylization pipeline from global scalars to per-location spatial maps, yielding local, per-axis control of content and style in a single generative pass.
- Because the two weights act on disjoint pathways, adjusting them independently spans a 2x2 retouching vocabulary, from free regeneration to identity preservation.
- We validate that edits stay confined to the retouched region and that each weight predominantly steers its own axis.
Sources (1)
- [1]Local Content-Style Control for Diffusion-based Image StylizationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:16 PM
Image stylization with latent-diffusion models entangles two independently refined axes: what a region depicts and how it is depicted.
Such pipelines expose only global controls, yet professional retouching demands deliberate, region-specific control.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Steering Diffusion Models to Rare Events with Sequential Monte Carlo
- Oct 6, 2026Enhancing Diffusion Language Models with Autoregressive Post-Training Weights
- Oct 6, 2026Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
- Oct 6, 2026Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size
- Oct 4, 2026How corner is a corner case? Percentile control for highway scenario generation
- Aug 10, 2026vllm-project/vllm v0.27.0