NTH

InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

AuthorsYunze Tong, Mushui Liu, Canyu Zhao, Shiyi Zhang, Didi Zhu, Peng Zhang, Wanggui He, Jinlong Liu, Ying Chen, Hao Jiang, Pipei Huang, Bo Zheng

August 24, 2026 3 min read
Watch on YouTube
The one-line take

InfinityEdit lets video generators apply ongoing edits to an endless stream while maintaining temporal continuity and stability over time.

Key results

14B
Helios backbone size

Frozen Helios-Distilled streaming video generator used as the backbone.

200
OOD benchmark videos

Source videos in the sequential-editing benchmark.

3
Editing rounds

Each benchmark sample contains three successive edit instructions.

0.7654
Camera Motion score

InfinityEdit's best VBench camera-motion score.

3.828
Edit Faithfulness score

Gemini-3.5-Flash judgment score on a 1–5 scale.

0.023
Cross-round faithfulness standard deviation

Edit-faithfulness variation across the three editing rounds.

What the paper found

InfinityEdit defines infinite video editing as generating each future video segment while applying successive natural-language edits, rather than rewriting a fixed clip frame by frame. The system attaches a lightweight Edit-Ignition Adapter to the frozen 14B Helios-Distilled streaming generator. Three attention modules provide complementary control: history cross-attention anchors the new chunk to prior frames, temporal causal self-attention propagates information forward, and edit cross-attention injects the instruction. The adapter activates only for the first chunk after an edit, while Helios continues subsequent chunks using a sliding history window and a reset anchor frame, preserving long-video generation with bounded memory. Training uses synthetic triplets from UltraVideo, detailed instructions expanded by Gemini 3 Flash, boundary-frame editing with Qwen-Image-Edit-2511, and target synthesis through Wan2.2-I2V-A14B, plus history corruption and a two-phase mixture-Gaussian flow-matching curriculum. On a benchmark of 200 source videos evaluated over 3 editing rounds, InfinityEdit achieved a 0.7654 Camera Motion score, compared with 0.5494 for the pure Helios baseline, and scored 3.828 for Edit Faithfulness under Gemini-3.5-Flash judging. Its edit-faithfulness standard deviation across rounds was only 0.023, indicating minimal degradation, and long-video tests remained stable beyond 1000 frames. The method targets open-ended use cases such as continuously restyling live footage or changing camera motion in an ongoing shot.

Original abstract

With large pretrained models, existing methods have effectively improved instruction-based video editing. However, most of them rely on an in-place editing assumption. They align the edited video with the given source clip frame by frame over a fixed time span. This pattern fails for open-ended streams, e.g., restyling a live game or applying a camera move to an ongoing shot. In such cases, edits must extend to future frames as they arrive, rather than be applied to a static input clip. In this paper, we study this setting and name it infinite video editing: given a preceding segment and an edit request, a model must generate the next segment that continues the stream while applying the requested edit. This process repeats as an unbounded sequence of edit instructions arrives. This task brings two challenges: the edit must be a faithful continuation rather than a frame-wise rewrite, and generation quality must remain stable as edits accumulate. To address them, we first design a data-collection pipeline for infinite video editing. Based on the collected data, we propose InfinityEdit, a lightweight edit adapter that equips a streaming video generator with unbounded editing ability. The adapter contains three attention modules. History cross-attention guides the denoising frames using the input frames. Temporal causal self-attention keeps temporal cues flowing only from earlier frames to later ones. Edit cross-attention injects the edit request into generation. During inference, the adapter is activated only in the chunk where an edit request arrives. Subsequent chunks are generated by the original model with a reset anchor frame. This scheme applies the edit while preserving the original model's infinite generation ability. Extensive experiments show that InfinityEdit faithfully continues the stream under each edit, and stays stable over unbounded edit sequences.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis