AION
Research paperMultimodal Models · Evaluation & Benchmarks · Image, Video & 3D Generation2 sources · Oct 8, 2026

Reasoning-Informed Visual Editing

To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.

Key points

  • Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but still face challenges in visual editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats.
  • We expand input formats to include multi-image conditioning and scale the benchmark to 1000 human-annotated test cases, released in English and Chinese.
  • Beyond benchmarking, we introduce RISE-Agent, a training-free agentic framework integrating reasoning-driven planning, tool-augmented execution, and verifier-guided refinement, outperforming most strong existing approaches across diverse RISE tasks.
  • We evaluate 58 visual editing approaches, including 34 open-source models, 19 closed-source models, and 5 agentic methods.

Sources (2)

  • [1]Reasoning-Informed Visual Editing
    Hugging Face Daily Papers · Oct 8, 12:00 AM
    To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.
    Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but still face challenges in visual editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats.
  • [2]Reasoning-Informed Visual Editing
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:15 PM · same content

Extractive summary: sentences quoted from the sources.