From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
AuthorsYitong Wang, Fangyun Wei, Jinjing Zhao, Sirui Zhang, Hongyang Zhang, Dong Chen, Bo Dai, Yan Lu
AffiliationsFudan University · Microsoft Research · The University of Sydney · USTC · University of Waterloo
Compo turns poster creation from describing a design in words into arranging and specifying its elements directly on a spatial canvas.
Key results
Approximate number of automatically constructed training samples.
Number of held-out benchmark samples.
Overall score on the binding-adherence benchmark.
Strongest general-purpose baseline’s overall benchmark score.
Poster-level accuracy requiring all bound text to render correctly.
What the paper found
This paper replaces prompt-only poster design with a Spatial Canvas: users place and specify elements directly in two-dimensional space using four bindings—semantic for concepts, identity for reference-preserved subjects, text for exact copy, and pixel for content to preserve. Its Compo model adapts pretrained image-editing transformers with LoRA, then uses DiffusionNFT reinforcement learning to optimize text accuracy, visual similarity, and layout adherence. An automated pipeline combines OpenAI’s GPT-5.5, Anthropic’s Claude Opus 4.6, and Google’s Gemini 3.5 Flash for planning, plus image-generation, segmentation, and OCR models, to create Compo-200K; evaluation uses a separate 500-sample benchmark. Compo also supports an agentic workflow in which GPT-5.6 turns a high-level request into a canvas plan. On the benchmark, Compo-LongCat reaches an overall average score of 0.752, compared with 0.504 for the strongest general-purpose baseline, HiDream-O1-Image, and achieves 0.940 sentence accuracy for poster text. The central contribution is a shared, explicit representation of layout, references, preserved pixels, and exact copy, letting generation focus on rendering a composed design rather than interpreting every constraint from prose.
Original abstract
Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words. We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance. Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications. Compo supports both direct inference, where users explicitly construct the canvas, and agentic mode, where a high-level request is automatically translated into a planned Spatial Canvas. To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch. We further introduce a benchmark that evaluates adherence to individual binding types and their joint composition. Experiments show that Compo achieves stronger compositional controllability than both general-purpose image generation models and dedicated poster generation systems while maintaining high visual quality. By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.
Read the original paperMore in Generative Models
Browse all 69 papers →Empirical Variational Autoencoder
Kaede Shiohara
EVA makes VAEs generate high-quality images and sounds faster by replacing their fixed Gaussian latent prior with a learned, self-predictive one.
MEND: RL For Flow Models via Proximal Velocity Matching
Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik
MEND makes flow-model reward tuning more efficient by moving samples only when the reward gain justifies the size of the change.
Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
Jiangshan Wang, Zeqiang Lai, Jiayi Guo, Xin Yang, Xin Huang, Jiarui Chen, Ziheng Ouyang, Chunchao Guo, Xiangyu Yue
Tex-Zero shows that high-quality native 3D textures may be learned from cleverly structured 2D images instead of costly real 3D assets.