NTH

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

AuthorsYueyi Liu, Chi Zhang, Sen Cui, Miao Liu

August 2, 2026 2 min read
Watch on YouTube
The one-line take

ElasticTTT keeps diffusion models from forgetting their generative skills while adapting them for one-shot video editing.

Key results

125
Evaluation test set

Text-video pairs covering five video-editing task types.

1.3B
Base model

Wan2.1 backbone used for the primary ElasticTTT experiments.

100
TTT optimization steps

Low-rank adaptation steps used during test-time tuning.

7%
Inference overhead

Additional overhead over vanilla TTT from Contrastive CFG and Async-NS.

6.68
Overall VLM score

ElasticTTT score under GPT-5 evaluation, compared with 5.39 for vanilla TTT.

7.00
Wan2.2-5B overall score

Overall score achieved by ElasticTTT on the larger Wan2.2-5B backbone.

What the paper found

Researchers Yueyi Liu, Chi Zhang, Sen Cui, and Miao Liu at Tsinghua University introduce ElasticTTT, a test-time tuning framework for diffusion-based video editing that addresses prior collapse: standard single-video optimization can cause conditioning collapse, where the model ignores target text, and spatial entanglement, where edited and preserved regions interfere. ElasticTTT combines Target Distribution Regularization, which injects zero-mean stochasticity into the optimization target without changing the expected flow-matching field; Contrastive Classifier-Free Guidance, which uses the source prompt as a negative constraint during sampling; and Asynchronous Noise Scheduling, which assigns separate noise trajectories to edited and preserved regions. Using the Wan2.1 1.3B model with 100 TTT steps, the method was evaluated on 125 text-video pairs spanning subject, background, addition, color, and general editing. With masks generated by Grounded-SAM2 and Qwen3-VL-32B, ElasticTTT achieved a GPT-5 evaluation overall score of 6.68, compared with 5.39 for vanilla TTT, while adding only 7% inference-time overhead on an NVIDIA GPU. On the larger Wan2.2-5B model, it reached 7.00 overall, versus 4.91 for baseline TTT. A human study scored ElasticTTT 3.60 overall, supporting its state-of-the-art balance between instruction adherence, visual quality, and source preservation.

Original abstract

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framework that preserves the prior generative distribution and rescues generative elasticity. Specifically, we propose \textit{Target Distribution Regularization} to prevent sharp memorization minima, \textit{Contrastive CFG} to guide inference away from source biases, and \textit{Asynchronous Noise Schedule} to preserve unedited regions. Extensive evaluations, supported by theoretical analysis, demonstrate that ElasticTTT successfully preserves the generative prior of the base model, achieving state-of-the-art performance on one-shot video editing.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis