StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation
AuthorsGuanlong Jiao, Chenyangguang Zhang, Jia Jun Cheng Xian, Zewei Zhang, Renjie Liao
Resources
This paper turns video editing into a fast, few-step generation problem, using streaming diffusion-style models and clever attention mechanisms to make edits both quicker and more controllable.
Key results
Best reported FiVE-Acc on FiVE-Bench for StreamGVE using LongLive with an edited first-frame visual prompt.
Text-driven StreamGVE built on Self Forcing reaches this FiVE-Acc on FiVE-Bench without visual prompting.
Text-driven StreamGVE built on LongLive reaches this FiVE-Acc on FiVE-Bench without visual prompting.
Reported inference speed on FiVE-Bench in 5-step settings, depending on the backbone and whether visual prompting is used.
StreamGVE is demonstrated on long videos exceeding 470 frames, showing length-unrestricted editing.
What the paper found
StreamGVE reframes training-free video editing as source-conditioned noise-to-data streaming generation, instead of the usual data-to-data inversion pipeline that needs many iterations and accumulates error. Built on pre-trained autoregressive streaming generators such as Self Forcing and LongLive, it performs dual-branch few-step sampling: the source video and target video are denoised in parallel under shared noise, then a self-attention bridge blends queries and keys to preserve structure, motion, and background while still allowing semantic change. For precise edit localization, StreamGVE derives masks from cross-attention grounding using foreground-background attention differences over trigger words, and then applies cross-attention boosting with a tunable strength ω to amplify edits only inside masked regions. It also adds source-oriented guidance, which corrects stochastic drift in editing-irrelevant regions by projecting source-branch velocity error onto the target branch, and an optional visual prompting mode that uses an edited first frame for finer control. On FiVE-Bench, StreamGVE achieves the best reported FiVE-Acc at 61.19 with LongLive plus visual prompting, while running at 0.60 to 0.76 seconds per frame in 5-step settings; without prompting, Ours (SF) reaches 51.94 FiVE-Acc and Ours (LL) reaches 55.04. The method generalizes to long videos over 470 frames and ranks first in user studies for quality, background consistency, and editing fidelity.
Original abstract
Although existing video editing methods are generally feasible, they often require many costly iterations and still struggle to deliver high-quality yet satisfying editing results. We attribute this limitation to the prevalent data-to-data paradigm, which is less compatible with modern generative models than noise-to-data generation. To address this gap, we revisit video editing from a noise-to-data perspective and propose Streaming-Generation-based Video Editing (StreamGVE), which preserves few-step sampling while seamlessly injecting source-video conditions. Built on pre-trained streaming generation models, StreamGVE introduces dual-branch fast sampling with a self-attention bridge and cross-attention grounding/boosting to satisfy both sampling and conditioning requirements. We further propose source-oriented guidance to improve target-generation quality, and a visual prompting strategy to enhance editing flexibility and practicality. The method is effective, robust, and generalizable across different models. Extensive experiments on diverse video editing tasks show that StreamGVE consistently outperforms existing approaches, even in few-step settings with minimal time cost.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.