One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
AuthorsAdheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen, Muntasir Wahed, Nabeel Bashir, Xiaona Zhou, Tianjiao Yu, Vedant Shah, Ismini Lourentzou
Resources
EditVid is a training-free video editor that combines several clever mechanisms to preserve motion, identity, and locality across many kinds of edits.
Key results
Overall edit-correctness score on FiVE for EditVid.
Strongest evaluated training-free baseline on FiVE.
Preference share over seven competing methods.
Result when sparse causal memory uses only the immediately preceding frame.
Seconds per frame on an NVIDIA H200.
What the paper found
EditVid is a training-free framework that turns a frozen image-editing multimodal diffusion transformer into a unified video editor for style transfer, attribute and part modification, object insertion, and reference-guided subject replacement. Built on FLUX.2-Klein-9B, with cross-backbone testing on FLUX.1-Kontext, it separates temporal consistency into three mechanisms: sparse causal memory reuses key–value states from the immediately preceding frame for local motion coherence; correspondence-based post-attention token injection transfers identity and appearance from an anchor frame over long ranges using confidence and cycle-consistency filtering; and soft latent blending preserves instruction-irrelevant regions without rigid masks. Using four denoising steps, EditVid reaches 78.16 FiVE-Acc, substantially above the strongest evaluated training-free baseline, FlowDirector, at 58.95, while remaining competitive on IVEBench and leading most criteria in a curated 50-video VLM evaluation. In a blind study with seven competing methods, it receives 51.8% overall preference. The best temporal-scope ablation uses only the immediately preceding frame, reaching 77.50 FiVE-Acc and 91.63 MF-S; retaining all previous frames instead degrades performance. On an NVIDIA H200, the system achieves 3.15 seconds per frame, demonstrating a practical accuracy–runtime trade-off without video-specific training or test-time optimization.
Original abstract
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8\% overall preference for EditVid over 7 competing methods.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.