NTH

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

AuthorsZhicong Tang, Zhao Zhang, Jingye Chen, Mohan Zhou, Yifan Pu, Yuchi Liu, Yalong Bai, Ethan Smith, Yuhui Yuan

June 10, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a large-scale diffusion model that can generate and edit editable layered images, making layered visual composition much faster and more practical.

Key results

20B
model size

MRT base architecture built on Qwen-Image

10M
training data size

curated multilingual layered design dataset

8
distilled denoising steps

few-step multi-layer generator after diffusion distillation

108.5x
speedup at ~20 layers

inference acceleration versus Qwen-Image-Layered

23.6x
memory reduction

peak GPU memory savings in image-to-layers inference

What the paper found

MRT, developed by Canva Research, is a 20B-parameter Masked Region Transformer for layered image generation and editing that unifies text-to-layers, image-to-layers, and layers-to-layers in one masked regional diffusion framework. Trained on over 10M multilingual design samples, the model operates on a full-size canvas with an overflow-aware background layer, so foreground layers can extend beyond the visible boundary instead of being cropped, which is crucial for editability and reuse. For inference, MRT applies diffusion distillation to compress a 50-step teacher into an 8-step generator, reaching near-real-time multi-layer synthesis while preserving quality. On image-to-layers, MRT beats Qwen-Image-Layered on merged-image quality and in user studies achieves 79.5%, 68.9%, and 82.6% win rates for layer quality, integrity, and granularity, while also delivering up to 108.5× speedup at about 20 layers and reducing peak GPU memory by 10.5× to 23.6×. The paper also shows that scaling matters: moving from a 13B FLUX.1-based baseline to Qwen-Image reduces FID from 17.79 to 16.15 on the same 0.5M subset, and expanding to the 10M dataset further lowers FID to 15.63. Overall, MRT establishes a new benchmark for editable, semi-transparent layered generation at scale.

Original abstract

Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editing in natural language. Despite its importance, this remains an underexplored area at scale. To address this gap, we present MRT, a 20B-parameter masked region diffusion model tailored for multi-layer transparent image generation and editing, trained on over 10M multilingual design samples spanning diverse aspect ratios and textual prompts. To fully leverage this scale, we make two key technical contributions. First, we unify three complementary tasks including text-to-layers, image-to-layers, and layers-to-layers within a shared masked region diffusion framework, where selective token masking enables flexible layer-wise generation and editing. Second, to enable overflow layer generation, we introduce an overflow-aware canvas layer that handles boundary inconsistencies and supports semi-transparent background synthesis, enabling complete editable layers extending beyond visible canvas boundaries. Additionally, we apply diffusion distillation to achieve 8-step, real-time multi-layer generation with minimal quality degradation. Extensive experiments demonstrate that our framework substantially outperforms prior state-of-the-art approaches, including various commercial systems, across all three tasks, establishing a new benchmark for multi-layer transparent image generation. Notably, our model significantly outperforms the concurrent Qwen-Image-Layered model in image-to-layers quality according to user-study results, while achieving 10-100\times faster inference and reducing activation GPU memory consumption by 50-90\% during image-to-layer inference.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis