Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation
To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
Key points
- Instruction-guided image editors have become highly capable, yet improving them further still depends on human-edited training pairs or external reward models.
- In this work, we strive to improve a pretrained image editor using only its own generations, without human-edited targets or an external training-time reward model.
- A Planner proposes structured edit instructions from unlabeled images, the Editor samples multiple candidate edits, and a frozen Critic scores each candidate with decomposed rubric checks for edit realization, removal of the old state, and content preservation, using features already exposed by the editor.
- On Qwen-Image-Edit, Rubric-CEPR improves ImgEdit from 4.36 to 4.60 (+5.5%), with a +24.9% gain on object isolation, and transfers to GEdit-Bench and Complex-Edit.
Sources (1)
- [1]Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-DistillationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:59 PM
To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
Instruction-guided image editors have become highly capable, yet improving them further still depends on human-edited training pairs or external reward models.
Extractive summary: sentences quoted from the sources.