AION
Research paperImage, Video & 3D Generation · Large Language Models · Multimodal Models2 sources · Oct 8, 2026

VibeEdit: Image Editing with Canvas Instructions

We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.

Key points

  • In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike.
  • We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing.
  • We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions.
  • We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects.

Sources (2)

  • [1]VibeEdit: Image Editing with Canvas Instructions
    Hugging Face Daily Papers · Oct 8, 12:00 AM
    We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.
    In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike.
  • [2]VibeEdit: Image Editing with Canvas Instructions
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:15 PM · same content

Extractive summary: sentences quoted from the sources.