NTH

EditaLive! Unified Character Video Editing for Live Streaming

AuthorsZhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun

September 2, 2026 2 min read
Watch on YouTube
The one-line take

EditaLive turns human video editing into a fast, live-streaming experience by adapting and distilling a pretrained animation model while preserving facial expressions and character appearance.

Key results

50K
CharEdit-50K dataset

Training dataset for appearance editing with motion-aligned source videos

2
Sampling steps

Distilled causal streaming sampler

0.492
Long-video identity similarity

ID-SIM on CharEdit-Bench-L

0.817
Long-video edit success rate

Success rate on CharEdit-Bench-L

14.47 FPS
Streaming throughput

End-to-end speed on one NVIDIA H100 GPU

0.829 seconds
Inter-chunk latency

Average streaming latency on one NVIDIA H100 GPU

What the paper found

EditaLive is a unified system for instruction-guided character video editing in live streams, built on Wan-Animate’s separation of appearance from motion. Instead of synthesizing paired edited videos that can distort facial expressions, it edits a reference frame and reconstructs the authentic source video using explicit body and facial-motion conditions, trained on the CharEdit-50K dataset. The framework then converts offline bidirectional generation into chunk-wise causal streaming and uses aligned self-rollout distillation, Fixed RoPE, and First-frame Preserved Sparse Attention to produce a stable 2-step sampler with reduced long-term appearance drift. CharEdit-50K combines instructions generated by OpenAI’s GPT-5.5 with edited references produced by Qwen-Image-Edit and Nano Banana 2. On CharEdit-Bench-L, EditaLive reaches an identity similarity of 0.492, facial-expression error of 0.576, edit success rate of 0.817, and 14.47 FPS with 0.829 seconds of inter-chunk latency on a single NVIDIA H100 GPU. The method also supports cross-character editing by combining one person’s appearance with another person’s motion, although single-image references can miss details in previously occluded regions and skeleton controls do not fully capture fine finger articulation.

Original abstract

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis