MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale
AuthorsZhicong Tang, Zhao Zhang, Jingye Chen, Mohan Zhou, Yifan Pu, Yuchi Liu, Yalong Bai, Ethan Smith, Yuhui Yuan
Resources
This paper introduces a large-scale diffusion model that can generate and edit editable layered images, making layered visual composition much faster and more practical.
Key results
MRT base architecture built on Qwen-Image
curated multilingual layered design dataset
few-step multi-layer generator after diffusion distillation
inference acceleration versus Qwen-Image-Layered
peak GPU memory savings in image-to-layers inference
What the paper found
MRT, developed by Canva Research, is a 20B-parameter Masked Region Transformer for layered image generation and editing that unifies text-to-layers, image-to-layers, and layers-to-layers in one masked regional diffusion framework. Trained on over 10M multilingual design samples, the model operates on a full-size canvas with an overflow-aware background layer, so foreground layers can extend beyond the visible boundary instead of being cropped, which is crucial for editability and reuse. For inference, MRT applies diffusion distillation to compress a 50-step teacher into an 8-step generator, reaching near-real-time multi-layer synthesis while preserving quality. On image-to-layers, MRT beats Qwen-Image-Layered on merged-image quality and in user studies achieves 79.5%, 68.9%, and 82.6% win rates for layer quality, integrity, and granularity, while also delivering up to 108.5× speedup at about 20 layers and reducing peak GPU memory by 10.5× to 23.6×. The paper also shows that scaling matters: moving from a 13B FLUX.1-based baseline to Qwen-Image reduces FID from 17.79 to 16.15 on the same 0.5M subset, and expanding to the 10M dataset further lowers FID to 15.63. Overall, MRT establishes a new benchmark for editable, semi-transparent layered generation at scale.
Original abstract
Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editing in natural language. Despite its importance, this remains an underexplored area at scale. To address this gap, we present MRT, a 20B-parameter masked region diffusion model tailored for multi-layer transparent image generation and editing, trained on over 10M multilingual design samples spanning diverse aspect ratios and textual prompts. To fully leverage this scale, we make two key technical contributions. First, we unify three complementary tasks including text-to-layers, image-to-layers, and layers-to-layers within a shared masked region diffusion framework, where selective token masking enables flexible layer-wise generation and editing. Second, to enable overflow layer generation, we introduce an overflow-aware canvas layer that handles boundary inconsistencies and supports semi-transparent background synthesis, enabling complete editable layers extending beyond visible canvas boundaries. Additionally, we apply diffusion distillation to achieve 8-step, real-time multi-layer generation with minimal quality degradation. Extensive experiments demonstrate that our framework substantially outperforms prior state-of-the-art approaches, including various commercial systems, across all three tasks, establishing a new benchmark for multi-layer transparent image generation. Notably, our model significantly outperforms the concurrent Qwen-Image-Layered model in image-to-layers quality according to user-study results, while achieving 10-100\times faster inference and reducing activation GPU memory consumption by 50-90\% during image-to-layer inference.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.