NTH

From Prompting to Composing: A Spatial Canvas Interface for Poster Generation

AuthorsYitong Wang, Fangyun Wei, Jinjing Zhao, Sirui Zhang, Hongyang Zhang, Dong Chen, Bo Dai, Yan Lu

AffiliationsFudan University · Microsoft Research · The University of Sydney · USTC · University of Waterloo

October 11, 2026 2 min read
Watch on YouTube
The one-line take

Compo turns poster creation from describing a design in words into arranging and specifying its elements directly on a spatial canvas.

Key results

200K
Compo-200K training dataset

Approximate number of automatically constructed training samples.

500
Benchmark size

Number of held-out benchmark samples.

0.752
Compo-LongCat overall average score

Overall score on the binding-adherence benchmark.

0.504
HiDream-O1-Image overall average score

Strongest general-purpose baseline’s overall benchmark score.

0.940
Compo-LongCat sentence accuracy

Poster-level accuracy requiring all bound text to render correctly.

What the paper found

This paper replaces prompt-only poster design with a Spatial Canvas: users place and specify elements directly in two-dimensional space using four bindings—semantic for concepts, identity for reference-preserved subjects, text for exact copy, and pixel for content to preserve. Its Compo model adapts pretrained image-editing transformers with LoRA, then uses DiffusionNFT reinforcement learning to optimize text accuracy, visual similarity, and layout adherence. An automated pipeline combines OpenAI’s GPT-5.5, Anthropic’s Claude Opus 4.6, and Google’s Gemini 3.5 Flash for planning, plus image-generation, segmentation, and OCR models, to create Compo-200K; evaluation uses a separate 500-sample benchmark. Compo also supports an agentic workflow in which GPT-5.6 turns a high-level request into a canvas plan. On the benchmark, Compo-LongCat reaches an overall average score of 0.752, compared with 0.504 for the strongest general-purpose baseline, HiDream-O1-Image, and achieves 0.940 sentence accuracy for poster text. The central contribution is a shared, explicit representation of layout, references, preserved pixels, and exact copy, letting generation focus on rendering a composed design rather than interpreting every constraint from prose.

Original abstract

Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words. We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance. Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications. Compo supports both direct inference, where users explicitly construct the canvas, and agentic mode, where a high-level request is automatically translated into a planned Spatial Canvas. To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch. We further introduce a benchmark that evaluates adherence to individual binding types and their joint composition. Experiments show that Compo achieves stronger compositional controllability than both general-purpose image generation models and dedicated poster generation systems while maintaining high visual quality. By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.

Read the original paper

More in Generative Models

Browse all 69 papers →
01Generative Model

Empirical Variational Autoencoder

Kaede Shiohara

EVA makes VAEs generate high-quality images and sounds faster by replacing their fixed Gaussian latent prior with a learned, self-predictive one.

Read analysis
02Generative Model

MEND: RL For Flow Models via Proximal Velocity Matching

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik

MEND makes flow-model reward tuning more efficient by moving samples only when the reward gain justifies the size of the change.

Read analysis