LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation
AuthorsQixin Hu, Shuai Yang, Wei Huang, Song Han, Yukang Chen
LongLive-RAG turns previously generated video latents into a searchable memory, helping autoregressive models keep long videos more consistent instead of drifting over time.
Key results
Retrieval embedding size used by the latent autoencoder
Milliseconds added per autoregressive block during a 120s rollout
Total retrieval time in milliseconds for a 120s rollout
Top-K retrieval setting that gave the strongest consistency and imaging quality
LongLive-RAG score on the 30s Causal-Forcing auxiliary video quality evaluation
What the paper found
LongLive-RAG, developed by NVIDIA researchers with affiliations including USC and MIT, reframes long video generation as retrieval-augmented generation over self-generated latents instead of relying only on sliding-window attention. The method keeps the base autoregressive video diffusion backbone frozen and adds a lightweight retrieval path: a query embedding from the latest completed latent searches a historical latent bank, then the top-K matched context entries are injected back into attention alongside the local window. To make retrieval useful rather than redundant, the paper trains a 1024-dimensional latent retrieval encoder with a reconstruction loss, a Window Temporal Delta Loss that suppresses near-duplicate short-range matches, and a second-order smoothing term to stabilize the embedding trajectory. Across three backbones—Causal-Forcing, Self-Forcing, and LongLive—and 30s, 60s, and 120s rollouts, LongLive-RAG achieves the best average VBench-Long rank in every setting, with especially strong gains in subject consistency, background consistency, motion smoothness, and imaging quality. The authors also report that the retrieval overhead is only 4.08 ms per block, or 490 ms total for a 120s rollout, and that K = 6 is the best retrieval budget under a fixed attention budget. On 30s Causal-Forcing, the full method improves auxiliary VLM exposure quality to 4.70 and beats random retrieval, average pooling, and reconstruction-only embeddings, supporting the claim that content-addressable memory is more effective than fixed anchors or compressed history for long-horizon video fidelity.
Original abstract
Autoregressive (AR) video diffusion enables variable-length synthesis, but long-horizon generation often suffers from accumulated errors and identity drift. For efficiency, existing methods commonly adopt sliding-window attention during generation. This creates an irreversible generation trajectory: once the active window accumulates appearance errors, subsequent generations can only condition on this degraded trajectory and drift further away. We address this limitation by formulating long video generation as a retrieval-augmented generation (RAG) problem. Rather than relying solely on the recent window, we treat previously generated latents as a dynamic, searchable history. We propose LongLive-RAG, a general retrieval framework for AR video generation. At each new block, LongLive-RAG uses a query embedding to retrieve relevant historical latents. This lightweight retrieval step adds only a small overhead relative to generation and lets the generator condition on non-local context instead of only the recent window. To make retrieval more discriminative, we introduce the Window Temporal Delta Loss that suppresses redundant local similarity and encourages embeddings to capture meaningful temporal changes. Together, these components help reduce error accumulation caused by sliding-window attention. Experiments across multiple AR backbones and generation lengths show improved long-video quality and the best average VBench-Long rank. To our knowledge, among open-ended AR long video generation methods, LongLive-RAG is the first to formulate self-generated latent history as content-addressable retrieval memory. Code is available at https://github.com/qixinhu11/LongLive-RAG.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.