AION
Research paperImage, Video & 3D Generation1 source · Oct 8, 2026

Transforming Image Editors into Video Editors

In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation.

Key points

  • Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly.
  • Our key insight is that video editing can be decomposed into two subproblems: editing a sparse set of keyframes and propagating those edits across time.
  • Based on this observation, we propose Anchor-based Video Editing (AVE), a two-stage framework in which a powerful image editor first performs composed editing on selected keyframes, and a motion-guided image-to-video diffusion model then generates the final video by treating the edited keyframes as fixed anchors.
  • Further ablations reveal that final video editing quality is strongly correlated with the quality of the image editor, suggesting that future progress in video editing may come from stronger image editing foundations and lightweight transfer to video.

Sources (1)

  • [1]Transforming Image Editors into Video Editors
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 12:43 AM
    In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation.
    Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly.

Extractive summary: sentences quoted from the sources.