ResearchResearch paperLarge Language Models · Image, Video & 3D Generation1 source · Oct 2, 2026

Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference.

Key points

  • In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals.
  • Our reward system consists of two main components: a preference reward, trained on large-scale human preference data using a Bradley-Terry objective to capture overall human aesthetic and perceptual preferences, and rubric-based rewards, which explicitly evaluate prompt faithfulness and other desirable properties while providing safeguards against reward hacking.
  • We show that a naive weighted average leads to suboptimal optimization behavior, and propose a simple reward composition strategy that more effectively balances preference optimization with rubric satisfaction.
  • To support reproducible research, we release Arena-T2I-Training, a 1K subset of training data that recovers some gains of full-scale training, providing a resource that we hope will facilitate future work on post-training for text-to-image models.

Sources (1)

  • [1]Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
    Hugging Face Daily Papers · Oct 2, 12:00 AM
    Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference.
    In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals.

Extractive summary: sentences quoted from the sources.

Related