Self-Consuming Generative Models with Co-Evolving Human Preferences
We show that when training relies entirely on user-curated synthetic data, iterative curation amplifies initial biases and drives the system toward one of multiple singleton equilibria in which the instance holding an initial advantage eventually dominates.
Key points
- Generative models are increasingly trained in self-consuming iterative loops, where users curate preferred samples from model-generated candidates and the curated samples are used to train future generations of the model.
- Prior work has largely assumed fixed user preferences, but in practice exposure to model outputs gradually reshapes what users perceive as desirable, creating a feedback loop in which model distributions and user preferences co-evolve.
- We take a first step toward understanding the long-term behavior of such coupled dynamics.
- Building on this insight, we study how reference-data injection can be used to control long-term outcomes, and propose an efficient algorithm that jointly selects a reference distribution and its mixing weight to steer the coupled system toward equilibria that preserve desired attributes while minimizing data collection costs.
Sources (1)
- [1]Self-Consuming Generative Models with Co-Evolving Human PreferencesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:18 AM
We show that when training relies entirely on user-curated synthetic data, iterative curation amplifies initial biases and drives the system toward one of multiple singleton equilibria in which the instance holding an initial advantage eventually dominates.
Generative models are increasingly trained in self-consuming iterative loops, where users curate preferred samples from model-generated candidates and the curated samples are used to train future generations of the model.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026World Models Dream of Success: Diagnosing and Repairing Failure Insensitivity in Robot World Models
- Oct 6, 2026Learning Transition Kernels of Jump-Diffusion Processes with Conditional Diffusion Models
- Oct 6, 2026Algorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny Transformers
- Oct 6, 2026Build a voice travel concierge with Amazon Bedrock AgentCore, Managed Knowledge Base and Nova Sonic
- Oct 6, 2026Improving Synthetic Data Generation for Argument Mining via Adversarial Reinforcement Learning
- Oct 4, 2026SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision
