NTH

Vera: A Layered Diffusion Model for Content-Preserving Video Editing

AuthorsHongkai Zheng, Ta-Ying Cheng, Benjamin Klein, Yisong Yue, Zhuoning Yuan

June 26, 2026 2 min read
Watch on YouTube
The one-line take

Vera edits videos by generating only the changed layer and blending it back with the original, aiming to preserve everything that should stay the same while still allowing strong visual edits.

Key results

486K
training frames

Layered training data used for Vera

25.3
object-addition PSNR

Vera-1.3B content preservation on object addition

35.2
background-change PSNR

Vera-1.3B content preservation on background change

26.1
object-addition PSNR

Vera-14B content preservation on object addition

36.2
background-change PSNR

Vera-14B content preservation on background change

513
valid user-study trials

Human 2AFC preference study trials

What the paper found

Vera, developed by California Institute of Technology and Netflix, Inc., is a layered diffusion framework for content-preserving video editing that avoids regenerating the full frame. Instead, it jointly predicts an edit layer, an alpha matte, and a natural composite video, then uses a Mixture-of-Transformers design with three separate DiTs and joint self-attention so the edit branch can stay creative while the preserved content is retained by construction. The method is trained with flow matching on a curated layered dataset of 486K frames at 832 × 480, built from synthetic composites and realistic single- and multi-object videos with accurate alpha mattes and effects such as shadows and reflections. On 72 object-addition tests and 69 background-change tests, Vera-1.3B reaches 25.3 dB PSNR for object addition and 35.2 dB for background change, while Vera-14B reaches 26.1 dB and 36.2 dB, respectively; it also tops instruction compliance on object addition with 4.13 IS and remains competitive on edit quality. The paper’s ablations show that the layered formulation, the composite branch, and MoT joint attention are all necessary, and a human 2AFC study over 513 valid trials preferred Vera on content preservation and instruction compliance.

Original abstract

Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservation remains a core challenge: existing methods regenerate every pixel and often alter elements that should remain unchanged, such as characters or background scenes. We introduce Vera, a layered diffusion framework for content-preserving video editing. Instead of regenerating the entire video, Vera generates an edit layer along with an alpha matte for compositing with the source video, separating creative editing from content preservation by design. To encourage coherent composition with the source video, we extend the text-to-video DiT into a Mixture-of-Transformers (MoT) architecture, with separate DiTs for each layer that interact through joint self-attention. To support the training of Vera, we further construct a high-quality layered dataset with accurate alpha mattes, diverse scenes and dynamics, and visual effects. Across our quantitative benchmark and human preference study, Vera outperforms leading open-source video editing models in content preservation while remaining competitive in edit quality, using 486K frames of layered training data.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis