From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.
Key points
- Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words.
- Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications.
- To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch.
- By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.
Sources (2)
- [1]From Prompting to Composing: A Spatial Canvas Interface for Poster GenerationHugging Face Daily Papers · Oct 8, 12:00 AM
We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.
Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words.
- [2]From Prompting to Composing: A Spatial Canvas Interface for Poster GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:16 PM · same content
Extractive summary: sentences quoted from the sources.