Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models
Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference.
Key points
- Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives.
- Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models.
- We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics.
- To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, enabling direct concept removal in FSD models while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities.
Sources (1)
- [1]Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:03 PM
Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference.
Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives.
Extractive summary: sentences quoted from the sources.