Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference.
ProofPaper ↗
Key points
- In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals.
- Our reward system consists of two main components: a preference reward, trained on large-scale human preference data using a Bradley-Terry objective to capture overall human aesthetic and perceptual preferences, and rubric-based rewards, which explicitly evaluate prompt faithfulness and other desirable properties while providing safeguards against reward hacking.
- We show that a naive weighted average leads to suboptimal optimization behavior, and propose a simple reward composition strategy that more effectively balances preference optimization with rubric satisfaction.
- To support reproducible research, we release Arena-T2I-Training, a 1K subset of training data that recovers some gains of full-scale training, providing a resource that we hope will facilitate future work on post-training for text-to-image models.
Sources (1)
- [1]Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric RewardsHugging Face Daily Papers · Oct 2, 12:00 AM
Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference.
In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals.
Extractive summary: sentences quoted from the sources.