NTH

FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring

AuthorsJiatong Li, Leo Liang, Linghe Kong, Yulun Zhang

August 2, 2026 2 min read
Watch on YouTube
The one-line take

FreqForcing uses low-frequency visual anchors to keep autoregressive video generation stable for minutes without additional training.

Key results

24
Long-horizon extrapolation

Extends 5-second training clips to 120-second generation, a 24-fold extrapolation.

59.58
VBench-Long Dynamic Degree at 60s

FreqForcing's Dynamic Degree score for 60-second videos.

20.94
VBench-Long Overall Consistency at 60s

FreqForcing's Overall Consistency score for 60-second videos.

16.5%
Inference latency increase

Additional latency over Self-Forcing on NVIDIA RTX A6000 GPUs.

1.3B
Base model size

Wan2.1-T2V-1.3B is used as the base model.

What the paper found

Researchers from Shanghai Jiao Tong University and Tencent HY Team propose FreqForcing, a training-free method for stabilizing autoregressive long-video diffusion. Unlike bidirectional systems such as Sora, HunyuanVideo, and Wan, autoregressive models repeatedly condition on their own imperfect outputs, causing color drift, motion stagnation, and collapse. The paper identifies this degradation as spectral energy drift concentrated in DC and low-frequency bands. Its Spectral Self-Anchoring mechanism uses two inference branches: local attention preserves high-frequency motion and detail, while anchor attention retrieves low-frequency structure from high-quality early frames; a 3D FFT-based low-pass fusion combines them, with temporal RoPE alignment maintaining valid positional ranges. Applied to the Wan2.1-T2V-1.3B base model, FreqForcing extends Self-Forcing trained on 5-second clips to stable 120-second generation, a 24-fold extrapolation, using the first 128 MovieGen prompts for evaluation. On VBench-Long at 60 seconds, it reaches 59.58 Dynamic Degree and 20.94 Overall Consistency, while at 120 seconds it records 58.97 and 20.98 respectively, outperforming training-free baselines and remaining competitive with training-based methods. The approach increases inference latency by only 16.5 percent on NVIDIA RTX A6000 GPUs, and ablations show that both attention sinks and spectral anchoring are necessary to suppress long-horizon drift.

Original abstract

Autoregressive video diffusion models enable real-time streaming video generation. However, errors introduced during self-rollout accumulate over long horizons, manifesting as color drift, motion stagnation, and eventual visual collapse. In this paper, we characterize this phenomenon from a frequency-domain perspective: error accumulation appears as a pronounced energy drift in the low-frequency bands. We further investigate the effectiveness of attention sink in the frequency domain, and find that it improves the video quality by alleviating the spectral energy drift to some extent, but cannot fully resolve it. Motivated by the above analysis, we propose FreqForcing, a training-free framework that addresses error accumulation in long-video generation via Spectral Self-Anchoring (SSA). The proposed SSA leverages the low-frequency components of anchor attention to maintain long-horizon visual stability, while preserving dynamic motion through the high-frequency components of local attention. Our FreqForcing extends Self-Forcing pretrained on 5s clips to two-minute generation, achieving 24x extrapolation. Extensive experiments show that FreqForcing outperforms existing training-free methods quantitatively and qualitatively while remaining competitive with representative training-based approaches.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis