Vera: A Layered Diffusion Model for Content-Preserving Video Editing
AuthorsHongkai Zheng, Ta-Ying Cheng, Benjamin Klein, Yisong Yue, Zhuoning Yuan
Resources
Vera edits videos by generating only the changed layer and blending it back with the original, aiming to preserve everything that should stay the same while still allowing strong visual edits.
Key results
Layered training data used for Vera
Vera-1.3B content preservation on object addition
Vera-1.3B content preservation on background change
Vera-14B content preservation on object addition
Vera-14B content preservation on background change
Human 2AFC preference study trials
What the paper found
Vera, developed by California Institute of Technology and Netflix, Inc., is a layered diffusion framework for content-preserving video editing that avoids regenerating the full frame. Instead, it jointly predicts an edit layer, an alpha matte, and a natural composite video, then uses a Mixture-of-Transformers design with three separate DiTs and joint self-attention so the edit branch can stay creative while the preserved content is retained by construction. The method is trained with flow matching on a curated layered dataset of 486K frames at 832 × 480, built from synthetic composites and realistic single- and multi-object videos with accurate alpha mattes and effects such as shadows and reflections. On 72 object-addition tests and 69 background-change tests, Vera-1.3B reaches 25.3 dB PSNR for object addition and 35.2 dB for background change, while Vera-14B reaches 26.1 dB and 36.2 dB, respectively; it also tops instruction compliance on object addition with 4.13 IS and remains competitive on edit quality. The paper’s ablations show that the layered formulation, the composite branch, and MoT joint attention are all necessary, and a human 2AFC study over 513 valid trials preferred Vera on content preservation and instruction compliance.
Original abstract
Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservation remains a core challenge: existing methods regenerate every pixel and often alter elements that should remain unchanged, such as characters or background scenes. We introduce Vera, a layered diffusion framework for content-preserving video editing. Instead of regenerating the entire video, Vera generates an edit layer along with an alpha matte for compositing with the source video, separating creative editing from content preservation by design. To encourage coherent composition with the source video, we extend the text-to-video DiT into a Mixture-of-Transformers (MoT) architecture, with separate DiTs for each layer that interact through joint self-attention. To support the training of Vera, we further construct a high-quality layered dataset with accurate alpha mattes, diverse scenes and dynamics, and visual effects. Across our quantitative benchmark and human preference study, Vera outperforms leading open-source video editing models in content preservation while remaining competitive in edit quality, using 486K frames of layered training data.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.