NTH

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

AuthorsAdheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen, Muntasir Wahed, Nabeel Bashir, Xiaona Zhou, Tianjiao Yu, Vedant Shah, Ismini Lourentzou

September 10, 2026 2 min read
Watch on YouTube
The one-line take

EditVid is a training-free video editor that combines several clever mechanisms to preserve motion, identity, and locality across many kinds of edits.

Key results

78.16
FiVE-Acc

Overall edit-correctness score on FiVE for EditVid.

58.95
FlowDirector FiVE-Acc

Strongest evaluated training-free baseline on FiVE.

51.8%
Overall user preference

Preference share over seven competing methods.

77.50
Best temporal-scope FiVE-Acc

Result when sparse causal memory uses only the immediately preceding frame.

3.15
Inference runtime

Seconds per frame on an NVIDIA H200.

What the paper found

EditVid is a training-free framework that turns a frozen image-editing multimodal diffusion transformer into a unified video editor for style transfer, attribute and part modification, object insertion, and reference-guided subject replacement. Built on FLUX.2-Klein-9B, with cross-backbone testing on FLUX.1-Kontext, it separates temporal consistency into three mechanisms: sparse causal memory reuses key–value states from the immediately preceding frame for local motion coherence; correspondence-based post-attention token injection transfers identity and appearance from an anchor frame over long ranges using confidence and cycle-consistency filtering; and soft latent blending preserves instruction-irrelevant regions without rigid masks. Using four denoising steps, EditVid reaches 78.16 FiVE-Acc, substantially above the strongest evaluated training-free baseline, FlowDirector, at 58.95, while remaining competitive on IVEBench and leading most criteria in a curated 50-video VLM evaluation. In a blind study with seven competing methods, it receives 51.8% overall preference. The best temporal-scope ablation uses only the immediately preceding frame, reaching 77.50 FiVE-Acc and 91.63 MF-S; retaining all previous frames instead degrades performance. On an NVIDIA H200, the system achieves 3.15 seconds per frame, demonstrating a practical accuracy–runtime trade-off without video-specific training or test-time optimization.

Original abstract

Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8\% overall preference for EditVid over 7 competing methods.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis