NTH

LongLive-Plug: Once-for-All Distillation for Video Generation

AuthorsShuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

AffiliationsNVIDIA

October 1, 2026 2 min read
Watch on YouTube
The one-line take

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Key results

54
Verified downstream models

Training-free deployment was verified across downstream video models.

3
Backbone families

Coverage included Wan2.1-14B, Wan2.2-TI2V-5B, and MiniMax-H3.

478.7
SCOPE transferred FVD

LongLive-Plug’s FVD after transfer with four-step sampling.

80
Shared distillation cost

Approximate H100 GPU-hours for reusable base-model distillation.

456.8
Task-specific cumulative cost

Approximate H100 GPU-hours for four separate task-specific distillation runs.

21%
Rank ablation transfer improvement

Transfer FVD improvement when LoRA rank increased from 16 to 128.

What the paper found

NVIDIA’s LongLive-Plug proposes once-for-all distillation for video generation: instead of retraining every specialized model, it distills three reusable LoRA capabilities on each backbone—a CFG adapter that replaces two-pass classifier-free guidance with one conditional evaluation, a DMD2 few-step adapter for four-step sampling, and a Streaming Long Tuning adapter for correcting errors in causal autoregressive long-video rollouts. The adapters preserve downstream task weights and transfer without target data or fine-tuning, including models with added conditioning branches or expanded output channels. Across 54 downstream models from three backbone families—Wan2.1-14B, Wan2.2-TI2V-5B, and MiniMax-H3—the method reduced SCOPE FVD from 805.5 with naive four-step sampling to 478.7, approaching task-specific distillation. Its decoupled CFG design makes adapter weight an approximate guidance dial while keeping the few-step adapter fixed, avoiding the collapse observed when scaling a coupled adapter. LongLive-Plug holds cumulative distillation at about 80 H100 GPU-hours, compared with 456.8 GPU-hours for four task-specific distillation runs. Transfer quality also depends on capacity and data: increasing LoRA rank from 16 to 128 improved transfer FVD by 21%, while less diverse prompts increased FVD by 12%. For long-context generation, the transferred adapter raised ReWorld’s seven-dimension VBench mean from 73.51 to 75.77 at 64 seconds and reached 84.34 on Matrix-Game 3.0 at 62.18 seconds, competitive with task-specific distillation.

Original abstract

Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis