AION
Research paperImage, Video & 3D Generation · Large Language Models · Robotics & Embodied AI2 sources · Oct 8, 2026

From Prompting to Composing: A Spatial Canvas Interface for Poster Generation

We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.

Key points

  • Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words.
  • Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications.
  • To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch.
  • By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.

Sources (2)

Extractive summary: sentences quoted from the sources.