ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
AuthorsYueyi Liu, Chi Zhang, Sen Cui, Miao Liu
Resources
ElasticTTT keeps diffusion models from forgetting their generative skills while adapting them for one-shot video editing.
Key results
Text-video pairs covering five video-editing task types.
Wan2.1 backbone used for the primary ElasticTTT experiments.
Low-rank adaptation steps used during test-time tuning.
Additional overhead over vanilla TTT from Contrastive CFG and Async-NS.
ElasticTTT score under GPT-5 evaluation, compared with 5.39 for vanilla TTT.
Overall score achieved by ElasticTTT on the larger Wan2.2-5B backbone.
What the paper found
Researchers Yueyi Liu, Chi Zhang, Sen Cui, and Miao Liu at Tsinghua University introduce ElasticTTT, a test-time tuning framework for diffusion-based video editing that addresses prior collapse: standard single-video optimization can cause conditioning collapse, where the model ignores target text, and spatial entanglement, where edited and preserved regions interfere. ElasticTTT combines Target Distribution Regularization, which injects zero-mean stochasticity into the optimization target without changing the expected flow-matching field; Contrastive Classifier-Free Guidance, which uses the source prompt as a negative constraint during sampling; and Asynchronous Noise Scheduling, which assigns separate noise trajectories to edited and preserved regions. Using the Wan2.1 1.3B model with 100 TTT steps, the method was evaluated on 125 text-video pairs spanning subject, background, addition, color, and general editing. With masks generated by Grounded-SAM2 and Qwen3-VL-32B, ElasticTTT achieved a GPT-5 evaluation overall score of 6.68, compared with 5.39 for vanilla TTT, while adding only 7% inference-time overhead on an NVIDIA GPU. On the larger Wan2.2-5B model, it reached 7.00 overall, versus 4.91 for baseline TTT. A human study scored ElasticTTT 3.60 overall, supporting its state-of-the-art balance between instruction adherence, visual quality, and source preservation.
Original abstract
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framework that preserves the prior generative distribution and rescues generative elasticity. Specifically, we propose \textit{Target Distribution Regularization} to prevent sharp memorization minima, \textit{Contrastive CFG} to guide inference away from source biases, and \textit{Asynchronous Noise Schedule} to preserve unedited regions. Extensive evaluations, supported by theoretical analysis, demonstrate that ElasticTTT successfully preserves the generative prior of the base model, achieving state-of-the-art performance on one-shot video editing.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.